CB-631: four credential env vars reach every member pane unblocked — the CB-596 policy has drifted #141

Closed
opened 2026-08-23 05:48:34 +02:00 by ltms · 2 comments
Owner

Found 2026-08-23 during the CB-623 (#127) worker-PR proof. Follows CB-596 and CB-608 (#111).

What the daemon reports on every single spawn

WARN HerdrPeerLauncher - memberCredentials gap: 4 credential-shaped env var name(s) are on
neither known: nor allow: — every member pane inherits them UNBLOCKED —
[AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH, N8N_WEBHOOK_TOKEN].
Add each to memberCredentials.known (blocked by default) or .allow (if a member legitimately needs it).

Two AWS keys, a Jenkins auth token and an n8n webhook token are inherited by every member pane, including members running on third-party backends.

This is a config gap, not a code defect

CB-596 made the credential policy config-driven and deny-by-default, and it is working exactly as designed: it noticed four names it does not know and said so, loudly, on every spawn. Nobody closed the gap. The policy in bridged/bridged.yaml lists 34 known / 7 allow / 29 blocked and has not kept up with what ${SHARED_ENV}/tools/secrets.sh now exports.

Note where the fix has to happen: bridged/bridged.yaml is gitignored. No worker can see it, no pull request will show the change, and CI cannot check it. This is the [gitignored-config-breaks-on-merge] shape — the lead has to make this edit directly and verify it on the live daemon.

Scope

  1. Decide per name: memberCredentials.known (blocked, the default and almost certainly right for all four) or .allow (only if a member legitimately needs it). None of these four has any reason to reach a member.
  2. Add them to bridged/bridged.yaml.
  3. Restart the daemon — memberCredentials is read per spawn, but confirm whether the surrounding config is snapshot-held before assuming a restart is unnecessary.
  4. Re-run scripts/probe-member-credentials.sh and check its denominator. CB-608 found that probe hardcodes its own list and under-reports; if it still does, this is a second instance of the same second-hand-list defect and should be fixed rather than worked around.

The wider question this raises

Four names appeared without anyone adding them to the policy. That means the gap will open again the next time secrets.sh grows. The warning fires on every spawn, which is good, but nothing fails and nothing tracks it, so it becomes noise.

Consider: should an unknown credential-shaped name be blocked by default rather than passed through with a warning? The current behaviour is fail-open on exactly the case the feature exists to prevent. If there is a reason it fails open, write it down in the config comment, because the next person will ask.

Acceptance criteria

  • A spawn logs no memberCredentials gap warning.
  • The probe script reports the number of names it checked alongside the number blocked, so a gap between the two is visible without knowing the second number.
  • The fail-open-vs-fail-closed decision is recorded in the config comment, whichever way it goes.
Found 2026-08-23 during the CB-623 (#127) worker-PR proof. Follows CB-596 and CB-608 (#111). ## What the daemon reports on every single spawn ``` WARN HerdrPeerLauncher - memberCredentials gap: 4 credential-shaped env var name(s) are on neither known: nor allow: — every member pane inherits them UNBLOCKED — [AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH, N8N_WEBHOOK_TOKEN]. Add each to memberCredentials.known (blocked by default) or .allow (if a member legitimately needs it). ``` Two AWS keys, a Jenkins auth token and an n8n webhook token are inherited by every member pane, including members running on third-party backends. ## This is a config gap, not a code defect CB-596 made the credential policy config-driven and deny-by-default, and it is working exactly as designed: it noticed four names it does not know and said so, loudly, on every spawn. Nobody closed the gap. The policy in `bridged/bridged.yaml` lists 34 known / 7 allow / 29 blocked and has not kept up with what `${SHARED_ENV}/tools/secrets.sh` now exports. **Note where the fix has to happen:** `bridged/bridged.yaml` is gitignored. No worker can see it, no pull request will show the change, and CI cannot check it. This is the [gitignored-config-breaks-on-merge] shape — the lead has to make this edit directly and verify it on the live daemon. ## Scope 1. Decide per name: `memberCredentials.known` (blocked, the default and almost certainly right for all four) or `.allow` (only if a member legitimately needs it). None of these four has any reason to reach a member. 2. Add them to `bridged/bridged.yaml`. 3. Restart the daemon — `memberCredentials` is read per spawn, but confirm whether the surrounding config is snapshot-held before assuming a restart is unnecessary. 4. Re-run `scripts/probe-member-credentials.sh` and check its denominator. CB-608 found that probe hardcodes its own list and under-reports; if it still does, this is a second instance of the same second-hand-list defect and should be fixed rather than worked around. ## The wider question this raises Four names appeared without anyone adding them to the policy. That means the gap will open again the next time `secrets.sh` grows. The warning fires on every spawn, which is good, but nothing fails and nothing tracks it, so it becomes noise. Consider: should an unknown credential-shaped name be **blocked by default** rather than passed through with a warning? The current behaviour is fail-open on exactly the case the feature exists to prevent. If there is a reason it fails open, write it down in the config comment, because the next person will ask. ## Acceptance criteria - A spawn logs no `memberCredentials gap` warning. - The probe script reports the number of names it checked alongside the number blocked, so a gap between the two is visible without knowing the second number. - The fail-open-vs-fail-closed decision is recorded in the config comment, whichever way it goes.
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-23 05:48:34 +02:00
Author
Owner

Correction after measuring this properly on 2026-08-23. One of the four names in this ticket
is not leaking, and the ticket's fix would not have closed the other three.
Details and the
replacement approach are in #144 (CB-633).

Checked each name three ways — exported by secrets.sh, present in its BRIDGED_MEMBER guard,
present in memberCredentials.known:

AWS_ACCESS_KEY_ID        exported_by_secrets=0  in_shell_guard=0  in_bridged.yaml=0
AWS_SECRET_ACCESS_KEY    exported_by_secrets=0  in_shell_guard=0  in_bridged.yaml=0
JENKINS_MCP_AUTH         exported_by_secrets=0  in_shell_guard=0  in_bridged.yaml=0
N8N_WEBHOOK_TOKEN        exported_by_secrets=1  in_shell_guard=1  in_bridged.yaml=0

Then simulated a member pane (BRIDGED_MEMBER=1 zsh -l, names and lengths only):

  • N8N_WEBHOOK_TOKEN -> BLOCKED-sentinel. It is already guarded. It is missing only from
    memberCredentials.known, which means the daemon cannot report on it — not that it leaks.
  • AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH -> leak in full.

They leak because they come from ${SHARED_ENV}/tools/mgnlSecrets.sh, a second secret file
with no guard at all
, which .ltms sources at line 41 — one line after the guarded
secrets.sh. So it also re-exports four names the guard had already blanked:
CONFLUENCE_USERNAME, CONFLUENCE_API_TOKEN, GITLAB_OAUTH_CLIENT_SECRET,
GITLAB_PERSONAL_ACCESS_TOKEN. That is 7 leaking credentials, not 4.

This ticket's fix was "add the four names to memberCredentials.known". That would have
changed nothing for the three real leaks, because the daemon-side overlay is overwritten by the
pane's login shell anyway — the shell guard is the only control that holds, and it is in a file
that runs before the file doing the leaking.

Keeping this ticket open only for the bookkeeping half: add N8N_WEBHOOK_TOKEN to
memberCredentials.known so the drift report can see it. The security half moves to #144.

Correction after measuring this properly on 2026-08-23. **One of the four names in this ticket is not leaking, and the ticket's fix would not have closed the other three.** Details and the replacement approach are in #144 (CB-633). Checked each name three ways — exported by `secrets.sh`, present in its `BRIDGED_MEMBER` guard, present in `memberCredentials.known`: ``` AWS_ACCESS_KEY_ID exported_by_secrets=0 in_shell_guard=0 in_bridged.yaml=0 AWS_SECRET_ACCESS_KEY exported_by_secrets=0 in_shell_guard=0 in_bridged.yaml=0 JENKINS_MCP_AUTH exported_by_secrets=0 in_shell_guard=0 in_bridged.yaml=0 N8N_WEBHOOK_TOKEN exported_by_secrets=1 in_shell_guard=1 in_bridged.yaml=0 ``` Then simulated a member pane (`BRIDGED_MEMBER=1 zsh -l`, names and lengths only): - `N8N_WEBHOOK_TOKEN` -> **BLOCKED-sentinel**. It is already guarded. It is missing only from `memberCredentials.known`, which means the daemon cannot *report* on it — not that it leaks. - `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `JENKINS_MCP_AUTH` -> **leak in full**. They leak because they come from `${SHARED_ENV}/tools/mgnlSecrets.sh`, a **second secret file with no guard at all**, which `.ltms` sources at line 41 — one line *after* the guarded `secrets.sh`. So it also re-exports four names the guard had already blanked: `CONFLUENCE_USERNAME`, `CONFLUENCE_API_TOKEN`, `GITLAB_OAUTH_CLIENT_SECRET`, `GITLAB_PERSONAL_ACCESS_TOKEN`. That is **7 leaking credentials**, not 4. This ticket's fix was "add the four names to `memberCredentials.known`". That would have changed nothing for the three real leaks, because the daemon-side overlay is overwritten by the pane's login shell anyway — the shell guard is the only control that holds, and it is in a file that runs before the file doing the leaking. Keeping this ticket open only for the bookkeeping half: add `N8N_WEBHOOK_TOKEN` to `memberCredentials.known` so the drift report can see it. The security half moves to #144.
Author
Owner

Done on the live daemon, 2026-09-03. All three acceptance criteria are met. This was a lead-side edit because fleetd/fleetd.yaml is gitignored — no worker can see it and no pull request can show it.

First, the ticket's premise has changed, and for the better

When this was filed the warning was a WARN and the four names really did reach every member pane unblocked. That is no longer true. policy: allow-list shipped on 2026-08-28 (#144). The scrub now keeps only names the allow-list derives, so an unknown name is blanked whether or not anyone listed it.

The log line says so itself now, and it dropped to INFO:

INFO d.l.f.m.HerdrPeerLauncher - memberCredentials gap: 5 credential-shaped env var name(s) are
on neither known: nor allow: — [AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH,
LLM_KEY_N8N, N8N_WEBHOOK_TOKEN]. The allow-list scrub blanks them anyway (they are not on the
derived allow-list), so no member pane keeps them; add each to memberCredentials.known...

So the leak this ticket was filed for is already closed, by a different change. The ticket's own question — "should an unknown credential-shaped name be blocked by default rather than passed through with a warning?" — has been answered yes, in code.

The count had also grown from 4 names to 5. LLM_KEY_N8N appeared after the ticket was written, which is exactly the drift the ticket predicted.

What I actually changed

Added all five to memberCredentials.known (blocked). None went to allow: — none of them has any reason to reach a member.

Live daemon, before and after, from GET /member-credentials:

before after
knownCount 34 39
blockedCount 29 34
allowedCount 7 7

No restart was needed. configReload is on at a 10-second interval and memberCredentials is read fresh per spawn, so the running daemon picked this up on its own. I checked that rather than assuming it.

Criterion 1 — a spawn logs no gap warning

Proven, not reasoned. I noted the exact log line number before the edit (44363), then spawned a fresh member, then counted gap lines written after that mark:

gap lines after the config change : 0
member credentials: allowed 16 of 82

Zero. The member was torn down afterwards.

Criterion 2 — the probe reports its denominator

Already true, and it has been since #111. scripts/probe-member-credentials.sh ends with:

$set_count of ${#NAMES[@]} names are set in this shell.
policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match.

It also refuses to run at all if the count it fetched and the count it parsed differ, or if either is zero. So the hardcoded-list defect this criterion was guarding against is gone, and no change was needed here. I did not edit the script.

Criterion 3 — the fail-open decision is written down

Recorded in the config comment above known:. The short version, for anyone reading the ticket rather than the file:

The list is fail-closed and has been since policy: allow-list. A gap is therefore no longer a leak. It is still worth closing, for two reasons. The per-spawn log line is the only thing that tells the operator the secret store has grown, and if the gap stays open that line becomes noise nobody reads. And if someone ever flips policy: back to deny-by-default, a stale list becomes a real leak again with no warning left to catch it.

The wider question the ticket raised

"Four names appeared without anyone adding them to the policy — the gap will open again the next time secrets.sh grows."

That is still true, and I am not fixing it here. The difference is that reopening the gap is now a housekeeping problem, not a security one: the scrub blocks the unknown name regardless. The log line is the tracker. It fires on every spawn, it names the exact names, and it tells you what to do. That seems proportionate now that it is not guarding a live leak.

Closing.

Done on the live daemon, 2026-09-03. All three acceptance criteria are met. This was a lead-side edit because `fleetd/fleetd.yaml` is gitignored — no worker can see it and no pull request can show it. ## First, the ticket's premise has changed, and for the better When this was filed the warning was a **WARN** and the four names really did reach every member pane unblocked. That is no longer true. `policy: allow-list` shipped on 2026-08-28 (#144). The scrub now keeps only names the allow-list derives, so an unknown name is blanked whether or not anyone listed it. The log line says so itself now, and it dropped to **INFO**: ``` INFO d.l.f.m.HerdrPeerLauncher - memberCredentials gap: 5 credential-shaped env var name(s) are on neither known: nor allow: — [AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH, LLM_KEY_N8N, N8N_WEBHOOK_TOKEN]. The allow-list scrub blanks them anyway (they are not on the derived allow-list), so no member pane keeps them; add each to memberCredentials.known... ``` So **the leak this ticket was filed for is already closed**, by a different change. The ticket's own question — "should an unknown credential-shaped name be blocked by default rather than passed through with a warning?" — has been answered yes, in code. The count had also grown from 4 names to 5. `LLM_KEY_N8N` appeared after the ticket was written, which is exactly the drift the ticket predicted. ## What I actually changed Added all five to `memberCredentials.known` (blocked). None went to `allow:` — none of them has any reason to reach a member. Live daemon, before and after, from `GET /member-credentials`: | | before | after | |---|---|---| | `knownCount` | 34 | **39** | | `blockedCount` | 29 | **34** | | `allowedCount` | 7 | 7 | No restart was needed. `configReload` is on at a 10-second interval and `memberCredentials` is read fresh per spawn, so the running daemon picked this up on its own. I checked that rather than assuming it. ## Criterion 1 — a spawn logs no gap warning Proven, not reasoned. I noted the exact log line number before the edit (44363), then spawned a fresh member, then counted gap lines written after that mark: ``` gap lines after the config change : 0 member credentials: allowed 16 of 82 ``` Zero. The member was torn down afterwards. ## Criterion 2 — the probe reports its denominator Already true, and it has been since #111. `scripts/probe-member-credentials.sh` ends with: ``` $set_count of ${#NAMES[@]} names are set in this shell. policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match. ``` It also refuses to run at all if the count it fetched and the count it parsed differ, or if either is zero. So the hardcoded-list defect this criterion was guarding against is gone, and no change was needed here. I did not edit the script. ## Criterion 3 — the fail-open decision is written down Recorded in the config comment above `known:`. The short version, for anyone reading the ticket rather than the file: The list is **fail-closed** and has been since `policy: allow-list`. A gap is therefore no longer a leak. It is still worth closing, for two reasons. The per-spawn log line is the only thing that tells the operator the secret store has grown, and if the gap stays open that line becomes noise nobody reads. And if someone ever flips `policy:` back to deny-by-default, a stale list becomes a real leak again with no warning left to catch it. ## The wider question the ticket raised "Four names appeared without anyone adding them to the policy — the gap will open again the next time `secrets.sh` grows." That is still true, and I am not fixing it here. The difference is that reopening the gap is now a housekeeping problem, not a security one: the scrub blocks the unknown name regardless. The log line is the tracker. It fires on every spawn, it names the exact names, and it tells you what to do. That seems proportionate now that it is not guarding a live leak. Closing.
ltms closed this issue 2026-09-03 11:25:08 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#141