CB-633: constrain a member's environment with an allow-list — the denylist misses a whole secret file #144

Open
opened 2026-08-23 06:20:51 +02:00 by ltms · 4 comments
Owner

Part of #125 (CB-621). Supersedes the approach in #141 (CB-631).

What is wrong

The fleet decides what credentials a member may see with a denylist: 29 names in
memberCredentials.known in bridged.yaml, mirrored by hand into a BRIDGED_MEMBER-guarded
block in ${SHARED_ENV}/tools/secrets.sh. The comment in that file says it plainly:

Keep this list in step with memberCredentials.known in bridged.yaml. They are two lists
that must agree.

Two hand-written lists that must agree is the defect family we already keep hitting (#113).
But the real problem is worse than drift. A denylist can only name what somebody remembered
to enumerate, and the exposure is not a file — it is an environment.

Measured, not assumed (2026-08-23)

${SHARED_ENV}/.ltms sources the secret files in this order:

line 40:  source ${SHARED_ENV}/tools/secrets.sh       # HAS the BRIDGED_MEMBER guard
line 41:  source ${SHARED_ENV}/tools/mgnlSecrets.sh    # has NO guard, and runs LAST

Nothing exports after line 41. Simulating a member pane exactly as herdr starts one:

$ BRIDGED_MEMBER=1 zsh -l -c '...'      # names + lengths only, never values
  AWS_ACCESS_KEY_ID                LEAKS (len=20)
  AWS_SECRET_ACCESS_KEY            LEAKS (len=40)
  JENKINS_MCP_AUTH                 LEAKS (len=56)
  CONFLUENCE_API_TOKEN             LEAKS (len=192)
  CONFLUENCE_USERNAME              LEAKS (len=23)
  GITLAB_OAUTH_CLIENT_SECRET       LEAKS (len=70)
  GITLAB_PERSONAL_ACCESS_TOKEN     LEAKS (len=51)
  GITEA_ACCESS_TOKEN               BLOCKED-sentinel
  N8N_WEBHOOK_TOKEN                BLOCKED-sentinel
  HASS_TOKEN                       BLOCKED-sentinel
  CF_API_TOKEN                     BLOCKED-sentinel

So 7 credentials reach every member pane in full. Four of them
(CONFLUENCE_*, GITLAB_*) are already named in the secrets.sh guard — the guard blanks
them, and then mgnlSecrets.sh exports them again one line later. The guard is beaten by
ordering, not by an incomplete list.

These are work credentials (Magnolia): a GitLab personal access token, a Confluence API token,
AWS keys, Jenkins auth. A member on a remote-model profile sends its context to a third-party
provider, so this is a real exposure and not a tidiness issue.

This also corrects #141: N8N_WEBHOOK_TOKEN is blocked, not leaking. The four names in
#141 were three genuine leaks plus one false positive.

Why the current design cannot be fixed by adding names

Adding AWS_* and JENKINS_MCP_AUTH to the two lists closes today's hole and nothing else.
The next secret file, or the next export in an existing one, reopens it silently. The
enumeration was of a file; the exposure is the environment (#113 family).

What to do instead

Invert it. In a member pane, blank everything that looks like a credential except an
explicit allow-list of what a member actually needs:

if [ -n "${FLEET_MEMBER:-}" ]; then
  _fleet_allow=" GITEA_TOKEN GITEA_HOST WORKER_GITEA_TOKEN CONTEXT7_TOKEN AI_GATEWAY_TOKEN "
  for _v in $(env | sed -n 's/^\([A-Za-z_][A-Za-z0-9_]*\)=.*/\1/p'); do
    case "$_v" in
      *TOKEN*|*KEY*|*SECRET*|*PASSWORD*|*AUTH*|*CREDENTIAL*)
        case "$_fleet_allow" in
          *" $_v "*) ;;
          *) export "$_v=blocked-by-fleet" ;;
        esac ;;
    esac
  done
  unset _v _fleet_allow
fi

Acceptance criteria:

  1. memberCredentials: in config gains policy: allow-list with an allow: list. known:
    becomes reporting-only — it stops being a control, so it stops being a thing to keep in sync.
  2. The allow-list is derived, not guessed. The launcher already knows every name it injects
    (gitTokenEnv -> GITEA_TOKEN, gitHostEnv -> GITEA_HOST, tokenEnv, and each profile's
    env: map). Build the allow-list from those, so adding a profile cannot break a spawn.
    Getting this wrong breaks every spawn, so it must not be a hand-typed list either.
  3. The daemon emits the shell block itself (a generated file the operator sources) instead of
    asking the operator to hand-mirror a list. One source, no drift.
  4. A test that starts a login shell with the marker set and asserts the allow-list survives
    and a sample of other credential-shaped names is blanked. Testing the launcher's env map is
    not enough — the launcher's map is what the login shell overwrites (#113 again: measure the
    side the failure is on).
  5. Report the denominator: log allowed N of M credential-shaped variables, not blocked N.

Stop-gap already staged

A BRIDGED_MEMBER-guarded block that blanks the 7 measured names, appended to the end of
.ltms so it runs after mgnlSecrets.sh. It is a denylist and is explicitly marked as a
stop-gap in its own comment. It is limited to names no member needs, so it cannot break a spawn.
Not yet applied — it edits a file outside this repo.

Part of #125 (CB-621). Supersedes the approach in #141 (CB-631). ## What is wrong The fleet decides what credentials a member may see with a **denylist**: 29 names in `memberCredentials.known` in `bridged.yaml`, mirrored by hand into a `BRIDGED_MEMBER`-guarded block in `${SHARED_ENV}/tools/secrets.sh`. The comment in that file says it plainly: > Keep this list in step with `memberCredentials.known` in `bridged.yaml`. They are two lists > that must agree. Two hand-written lists that must agree is the defect family we already keep hitting (#113). But the real problem is worse than drift. **A denylist can only name what somebody remembered to enumerate, and the exposure is not a file — it is an environment.** ## Measured, not assumed (2026-08-23) `${SHARED_ENV}/.ltms` sources the secret files in this order: ``` line 40: source ${SHARED_ENV}/tools/secrets.sh # HAS the BRIDGED_MEMBER guard line 41: source ${SHARED_ENV}/tools/mgnlSecrets.sh # has NO guard, and runs LAST ``` Nothing exports after line 41. Simulating a member pane exactly as herdr starts one: ``` $ BRIDGED_MEMBER=1 zsh -l -c '...' # names + lengths only, never values AWS_ACCESS_KEY_ID LEAKS (len=20) AWS_SECRET_ACCESS_KEY LEAKS (len=40) JENKINS_MCP_AUTH LEAKS (len=56) CONFLUENCE_API_TOKEN LEAKS (len=192) CONFLUENCE_USERNAME LEAKS (len=23) GITLAB_OAUTH_CLIENT_SECRET LEAKS (len=70) GITLAB_PERSONAL_ACCESS_TOKEN LEAKS (len=51) GITEA_ACCESS_TOKEN BLOCKED-sentinel N8N_WEBHOOK_TOKEN BLOCKED-sentinel HASS_TOKEN BLOCKED-sentinel CF_API_TOKEN BLOCKED-sentinel ``` So **7 credentials reach every member pane in full**. Four of them (`CONFLUENCE_*`, `GITLAB_*`) are already named in the `secrets.sh` guard — the guard blanks them, and then `mgnlSecrets.sh` exports them again one line later. The guard is beaten by ordering, not by an incomplete list. These are work credentials (Magnolia): a GitLab personal access token, a Confluence API token, AWS keys, Jenkins auth. A member on a remote-model profile sends its context to a third-party provider, so this is a real exposure and not a tidiness issue. This also corrects #141: `N8N_WEBHOOK_TOKEN` is **blocked**, not leaking. The four names in #141 were three genuine leaks plus one false positive. ## Why the current design cannot be fixed by adding names Adding `AWS_*` and `JENKINS_MCP_AUTH` to the two lists closes today's hole and nothing else. The next secret file, or the next `export` in an existing one, reopens it silently. The enumeration was of a file; the exposure is the environment (#113 family). ## What to do instead Invert it. In a member pane, blank **everything** that looks like a credential except an explicit allow-list of what a member actually needs: ```sh if [ -n "${FLEET_MEMBER:-}" ]; then _fleet_allow=" GITEA_TOKEN GITEA_HOST WORKER_GITEA_TOKEN CONTEXT7_TOKEN AI_GATEWAY_TOKEN " for _v in $(env | sed -n 's/^\([A-Za-z_][A-Za-z0-9_]*\)=.*/\1/p'); do case "$_v" in *TOKEN*|*KEY*|*SECRET*|*PASSWORD*|*AUTH*|*CREDENTIAL*) case "$_fleet_allow" in *" $_v "*) ;; *) export "$_v=blocked-by-fleet" ;; esac ;; esac done unset _v _fleet_allow fi ``` Acceptance criteria: 1. `memberCredentials:` in config gains `policy: allow-list` with an `allow:` list. `known:` becomes reporting-only — it stops being a control, so it stops being a thing to keep in sync. 2. **The allow-list is derived, not guessed.** The launcher already knows every name it injects (`gitTokenEnv` -> `GITEA_TOKEN`, `gitHostEnv` -> `GITEA_HOST`, `tokenEnv`, and each profile's `env:` map). Build the allow-list from those, so adding a profile cannot break a spawn. Getting this wrong breaks every spawn, so it must not be a hand-typed list either. 3. The daemon emits the shell block itself (a generated file the operator sources) instead of asking the operator to hand-mirror a list. One source, no drift. 4. A test that starts a **login shell** with the marker set and asserts the allow-list survives and a sample of other credential-shaped names is blanked. Testing the launcher's env map is not enough — the launcher's map is what the login shell overwrites (#113 again: measure the side the failure is on). 5. Report the denominator: log `allowed N of M credential-shaped variables`, not `blocked N`. ## Stop-gap already staged A `BRIDGED_MEMBER`-guarded block that blanks the 7 measured names, appended to the **end** of `.ltms` so it runs after `mgnlSecrets.sh`. It is a denylist and is explicitly marked as a stop-gap in its own comment. It is limited to names no member needs, so it cannot break a spawn. Not yet applied — it edits a file outside this repo.
Author
Owner

The vms lead reproduced this independently and found two things that change the ticket.

1. The argument in this ticket was too weak

I wrote that "a denylist can only name what somebody remembered to enumerate". That is true but
it is not what happened here. Four of the seven leaking names are already on the guard's
list.
The guard ran, blanked them correctly, and then .ltms line 41 sourced
mgnlSecrets.sh, which handed the real values straight back.

So a correct and complete denylist still failed. The defect is not an incomplete list — it is
that the control lives inside a sourced file, and a file sourced later can always undo it.
That is host-independent: it would happen the same way on a fresh vhost.

This sharpens acceptance criterion 2. The allow-list must be applied at the spawn boundary,
where the launcher builds the child environment and nothing runs after it. An allow-list that is
itself a block inside a sourced shell file inherits the exact defect it is meant to fix.

Independent confirmation, by comparing SHA-256 prefixes of each value between a normal login
shell and a BRIDGED_MEMBER=1 login shell (prefixes only, never values):

                                normal        member
N8N_WEBHOOK_TOKEN               dd92c8769933  fe8f1af3c3f9   BLOCKED
GITEA_ACCESS_TOKEN              bb8495c75917  fe8f1af3c3f9   BLOCKED
TELEGRAM_BOT_TOKEN              720f23c18a39  fe8f1af3c3f9   BLOCKED
AWS_ACCESS_KEY_ID               60f27c17e7c8  60f27c17e7c8   LEAKS
AWS_SECRET_ACCESS_KEY           b9710cda4901  b9710cda4901   LEAKS
JENKINS_MCP_AUTH                187822fb074e  187822fb074e   LEAKS
CONFLUENCE_USERNAME             2ad5945804f5  2ad5945804f5   LEAKS
CONFLUENCE_API_TOKEN            0f3153f2c245  0f3153f2c245   LEAKS
GITLAB_OAUTH_CLIENT_SECRET      9cb2e5d471f2  9cb2e5d471f2   LEAKS
GITLAB_PERSONAL_ACCESS_TOKEN    44660124b18c  44660124b18c   LEAKS

Three different credentials collapsing to the same hash prefix is what a working guard looks
like. Comparing hashes is a better probe than the sentinel-string check I used, because it needs
no knowledge of what the sentinel is.

2. The AWS key is not a limited one — raise the priority

AWS_ACCESS_KEY_ID carries AdministratorAccess via the admins group on account
891377113284. Created 2024-03-20, never rotated, last used three days ago.

That makes this the most serious of the seven by a wide margin, and it changes the ordering:
rotate the AWS key first, ahead of the other six and ahead of building anything here. A
full-admin, never-rotated key reached every member pane we have ever spawned on a remote-model
profile.

Rotation is the operator's call and is not part of this ticket. Noting it here so the ticket
records the real weight of what leaked, rather than treating all seven as equivalent.

3. The vhost does not replace this

Considered and rejected. Members on a separate host do not inherit the operator's Mac login
environment, so the vhost shrinks the blast radius — but it does not fix the defect, for the
reason in section 1: source ordering is host-independent. It also does not help while the lead
still runs on the Mac. Build the allow-list; treat the vhost as a separate win.

The `vms` lead reproduced this independently and found two things that change the ticket. ## 1. The argument in this ticket was too weak I wrote that "a denylist can only name what somebody remembered to enumerate". That is true but it is not what happened here. **Four of the seven leaking names are already on the guard's list.** The guard ran, blanked them correctly, and then `.ltms` line 41 sourced `mgnlSecrets.sh`, which handed the real values straight back. So a *correct and complete* denylist still failed. The defect is not an incomplete list — it is that the control lives **inside a sourced file**, and a file sourced later can always undo it. That is host-independent: it would happen the same way on a fresh vhost. This sharpens acceptance criterion 2. The allow-list must be applied at the **spawn boundary**, where the launcher builds the child environment and nothing runs after it. An allow-list that is itself a block inside a sourced shell file inherits the exact defect it is meant to fix. Independent confirmation, by comparing SHA-256 prefixes of each value between a normal login shell and a `BRIDGED_MEMBER=1` login shell (prefixes only, never values): ``` normal member N8N_WEBHOOK_TOKEN dd92c8769933 fe8f1af3c3f9 BLOCKED GITEA_ACCESS_TOKEN bb8495c75917 fe8f1af3c3f9 BLOCKED TELEGRAM_BOT_TOKEN 720f23c18a39 fe8f1af3c3f9 BLOCKED AWS_ACCESS_KEY_ID 60f27c17e7c8 60f27c17e7c8 LEAKS AWS_SECRET_ACCESS_KEY b9710cda4901 b9710cda4901 LEAKS JENKINS_MCP_AUTH 187822fb074e 187822fb074e LEAKS CONFLUENCE_USERNAME 2ad5945804f5 2ad5945804f5 LEAKS CONFLUENCE_API_TOKEN 0f3153f2c245 0f3153f2c245 LEAKS GITLAB_OAUTH_CLIENT_SECRET 9cb2e5d471f2 9cb2e5d471f2 LEAKS GITLAB_PERSONAL_ACCESS_TOKEN 44660124b18c 44660124b18c LEAKS ``` Three different credentials collapsing to the same hash prefix is what a working guard looks like. Comparing hashes is a better probe than the sentinel-string check I used, because it needs no knowledge of what the sentinel is. ## 2. The AWS key is not a limited one — raise the priority `AWS_ACCESS_KEY_ID` carries **`AdministratorAccess`** via the `admins` group on account `891377113284`. Created 2024-03-20, never rotated, last used three days ago. That makes this the most serious of the seven by a wide margin, and it changes the ordering: **rotate the AWS key first**, ahead of the other six and ahead of building anything here. A full-admin, never-rotated key reached every member pane we have ever spawned on a remote-model profile. Rotation is the operator's call and is not part of this ticket. Noting it here so the ticket records the real weight of what leaked, rather than treating all seven as equivalent. ## 3. The vhost does not replace this Considered and rejected. Members on a separate host do not inherit the operator's Mac login environment, so the vhost shrinks the blast radius — but it does not fix the defect, for the reason in section 1: source ordering is host-independent. It also does not help while the lead still runs on the Mac. Build the allow-list; treat the vhost as a separate win.
Author
Owner

Stop-gap applied by the operator, and verified here independently

The operator fixed it a better way than my staged block: reorder the two sources so
mgnlSecrets.sh runs before the guarded secrets.sh, plus add the three missing names to the
guard list. My block was not applied and is not needed.

Checked myself rather than taking the report — .ltms now reads:

line 40:  source ${SHARED_ENV}/tools/mgnlSecrets.sh
line 41:  source ${SHARED_ENV}/tools/secrets.sh      # guard now runs LAST

and a member login shell started from a clean parent (env -i HOME=... BRIDGED_MEMBER=1 zsh -l, so nothing is inherited from my own session):

AWS_ACCESS_KEY_ID              BLOCKED     CONFLUENCE_API_TOKEN          BLOCKED
AWS_SECRET_ACCESS_KEY          BLOCKED     GITLAB_OAUTH_CLIENT_SECRET    BLOCKED
JENKINS_MCP_AUTH               BLOCKED     GITLAB_PERSONAL_ACCESS_TOKEN  BLOCKED
CONFLUENCE_USERNAME            BLOCKED

A full sweep of every credential-shaped name in that shell leaves exactly the three intended
pass-throughs and nothing else:

SURVIVES: AI_GATEWAY_TOKEN (48)   CONTEXT7_TOKEN (64)   WORKER_GITEA_TOKEN (40)

The reorder is safe because mgnlSecrets.sh contains no variable references, so nothing in it
depended on secrets.sh having run first. Confirmed by reading it.

This ticket stays open

The exposure is closed. The defect is not fixed, and the fix makes that clearer rather than
less clear: the control is still a block inside a sourced file, and it now works only because
of the order two source lines happen to be in. Add a third secret file after line 41 — or move
one line — and it reopens silently, with no test anywhere that would notice.

That is the same shape as the original bug. The reorder buys time; it is not the answer.
Acceptance criteria unchanged: apply the allow-list at the spawn boundary, in the launcher,
where nothing runs afterwards.

One extra criterion this episode earns:

  1. A test that would have caught the reopening. Start a member login shell from a clean
    parent and assert the surviving set equals the allow-list exactly — an equality assertion,
    not a "these names are blocked" assertion. A blocked-name list is the very thing that fails
    here, because it can only check names somebody already thought of.
## Stop-gap applied by the operator, and verified here independently The operator fixed it a better way than my staged block: **reorder the two sources** so `mgnlSecrets.sh` runs *before* the guarded `secrets.sh`, plus add the three missing names to the guard list. My block was not applied and is not needed. Checked myself rather than taking the report — `.ltms` now reads: ``` line 40: source ${SHARED_ENV}/tools/mgnlSecrets.sh line 41: source ${SHARED_ENV}/tools/secrets.sh # guard now runs LAST ``` and a member login shell started from a **clean parent** (`env -i HOME=... BRIDGED_MEMBER=1 zsh -l`, so nothing is inherited from my own session): ``` AWS_ACCESS_KEY_ID BLOCKED CONFLUENCE_API_TOKEN BLOCKED AWS_SECRET_ACCESS_KEY BLOCKED GITLAB_OAUTH_CLIENT_SECRET BLOCKED JENKINS_MCP_AUTH BLOCKED GITLAB_PERSONAL_ACCESS_TOKEN BLOCKED CONFLUENCE_USERNAME BLOCKED ``` A full sweep of every credential-shaped name in that shell leaves exactly the three intended pass-throughs and nothing else: ``` SURVIVES: AI_GATEWAY_TOKEN (48) CONTEXT7_TOKEN (64) WORKER_GITEA_TOKEN (40) ``` The reorder is safe because `mgnlSecrets.sh` contains no variable references, so nothing in it depended on `secrets.sh` having run first. Confirmed by reading it. ## This ticket stays open The exposure is closed. The **defect is not fixed**, and the fix makes that clearer rather than less clear: the control is still a block inside a sourced file, and it now works only because of the order two `source` lines happen to be in. Add a third secret file after line 41 — or move one line — and it reopens silently, with no test anywhere that would notice. That is the same shape as the original bug. The reorder buys time; it is not the answer. Acceptance criteria unchanged: apply the allow-list at the **spawn boundary**, in the launcher, where nothing runs afterwards. One extra criterion this episode earns: 6. **A test that would have caught the reopening.** Start a member login shell from a clean parent and assert the surviving set equals the allow-list exactly — an equality assertion, not a "these names are blocked" assertion. A blocked-name list is the very thing that fails here, because it can only check names somebody already thought of.
Author
Owner

The seam exists after all — ZDOTDIR. Measured on this host, not reasoned about.

This overturns the conclusion in 7930a31 ("CB-596 round 2: the exec-time argv-prefix fix has no seam — stop and report"). That commit was right about what it checked and wrong about what it concluded. It checked herdr's protocol: agent.start takes a fixed kind (herdr resolves the executable) plus trailing args, and only tab.create/pane.split carry an env map. All still true — I re-read AgentControl.start and WorkspaceControl.createTab today.

But "no control point in the herdr protocol" is not "no control point". The shell itself has one, and we already reach it: tab.create's env map runs before the shell starts, and zsh reads its startup files from $ZDOTDIR. So the daemon can redirect where the member's shell looks for its own rc files, and put its scrub in the file that runs last.

Why this beats every earlier attempt

Startup order for a login interactive zsh is .zshenv → .zprofile → .zshrc → .zlogin. The operator's chain (.zshrc → .ltms → mgnlSecrets.sh + secrets.sh) all happens inside the first three. .zlogin runs after all of it. Nothing the operator sources can undo it, and the fix depends on no file outside this repo — which is the exact property the reorder stop-gap lacks.

Measured

ZDOTDIR is honoured by a login zsh started from a clean parent:

$ env -i HOME=$HOME PATH=/usr/bin:/bin ZDOTDIR=$S /bin/zsh -l -c '...'
MARK .zshenv   ZDOTDIR=/…/zdot
MARK .zprofile ZDOTDIR=/…/zdot
MARK .zlogin   ZDOTDIR=/…/zdot

Then the full design: four generated files in $ZDOTDIR, each sourcing the operator's real $HOME/.z* counterpart, with the scrub appended to .zlogin. Run as an interactive login shell from a clean parent (env -i HOME=… ZDOTDIR=… /bin/zsh -l -i), so nothing is inherited from my session:

fleet: allowed 3 of 28 credential-shaped variables

SURVIVES  AI_GATEWAY_TOKEN (48)   CONTEXT7_TOKEN (64)   WORKER_GITEA_TOKEN (40)

BLOCKED   AWS_ACCESS_KEY_ID · AWS_SECRET_ACCESS_KEY · JENKINS_MCP_AUTH ·
          CONFLUENCE_API_TOKEN · GITLAB_PERSONAL_ACCESS_TOKEN · GITEA_ACCESS_TOKEN ·
          TELEGRAM_BOT_TOKEN · N8N_WEBHOOK_TOKEN · N8N_ENCRYPTION_KEY ·
          N8N_OWNER_PASSWORD · HASS_TOKEN · CF_API_TOKEN · CF_USER_TOKEN ·
          TS_AUTHKEY · TS_API_KEY · BESZEL_ADMIN_PASSWORD · BESZEL_KEY ·
          BESZEL_UNIVERSAL_TOKEN · GRAFANA_ADMIN_PASSWORD · METRICS_PUSH_TOKEN ·
          BRAIN_MCP_TOKEN · MEMORY_MCP_TOKEN · LTMS_API_KEY · HW_PASSWORD · …

That is the AWS AdministratorAccess key blocked by the daemon's own file, with the operator's .ltms reordering removed from the equation entirely.

I also checked that no operator rc file sets ZDOTDIR (grep across .zshenv .zprofile .zshrc .zlogin .zlogout → no match), so the daemon's value is not overwritten.

The one thing this measurement exposes about the pattern list

CONFLUENCE_USERNAME was not blocked. It matches none of *TOKEN* *KEY* *SECRET* *PASSWORD* *AUTH* *CREDENTIAL*, because a username is the other half of a credential and is not shaped like one. That is this ticket's own defect in miniature: a pattern list is an enumeration, and it can only catch what somebody's pattern happened to describe.

So the policy should be a real allow-list: blank every variable that is neither on the derived allow-list nor on a small infrastructure passthrough set (PATH, HOME, SHELL, TERM, LANG, TMPDIR, USER, PWD, SSH_AUTH_SOCK² …). Then a credential nobody has thought of is blocked because it is new, not because it matched. Credential-shaped patterns stay only as a warning signal in the log, never as the control.

² SSH_AUTH_SOCK is #110 (CB-607) — it is a handle to the operator's agent, not a value in any secret file, which is why three credential tickets missed it. Decide it explicitly here rather than letting it pass through by silence.

Revised acceptance criteria

  1. memberCredentials: gains policy: allow-list. known: becomes reporting-only.
  2. The allow-list is derived from what the launcher itself injects — gitTokenEnv, gitHostEnv, tokenEnv, and every profile's env: map — plus the infrastructure passthrough set. Not hand-typed: a hand-typed list breaks every spawn the first time a profile adds a variable.
  3. The daemon generates the $ZDOTDIR directory per spawn and passes ZDOTDIR in tab.create's env map. No operator-owned file is edited, and nothing needs keeping in sync. Each generated file sources its $HOME counterpart first so PATH and the agent binaries still resolve.
  4. Non-zsh shells: if the member's shell is not zsh, ZDOTDIR does nothing and the member is unprotected. Detect it and refuse the spawn, or log a loud WARN — do not fail silently. Say which was chosen and why.
  5. A test that starts a real login shell from a clean parent and asserts the surviving set equals the allow-list. Equality, not "these names are blocked" — a blocked-name list is the very thing that fails here.
  6. Log the denominator: allowed N of M, never blocked N (#113).
  7. One live spawn as proof. The artefact is a shell file handed to another program; a unit test on the generator proves nothing about zsh (#113 instance 4).

What this does not change

The .ltms reorder stays as defence in depth. Rotating the AWS key is still first and still the operator's call.

## The seam exists after all — `ZDOTDIR`. Measured on this host, not reasoned about. This overturns the conclusion in `7930a31` ("CB-596 round 2: the exec-time argv-prefix fix has no seam — stop and report"). That commit was right about what it checked and wrong about what it concluded. It checked herdr's **protocol**: `agent.start` takes a fixed `kind` (herdr resolves the executable) plus trailing args, and only `tab.create`/`pane.split` carry an `env` map. All still true — I re-read `AgentControl.start` and `WorkspaceControl.createTab` today. But "no control point in the herdr protocol" is not "no control point". The shell itself has one, and we already reach it: **`tab.create`'s env map runs before the shell starts, and zsh reads its startup files from `$ZDOTDIR`.** So the daemon can redirect where the member's shell looks for its own rc files, and put its scrub in the file that runs *last*. ### Why this beats every earlier attempt Startup order for a login interactive zsh is `.zshenv` → `.zprofile` → `.zshrc` → **`.zlogin`**. The operator's chain (`.zshrc` → `.ltms` → `mgnlSecrets.sh` + `secrets.sh`) all happens inside the first three. `.zlogin` runs after all of it. Nothing the operator sources can undo it, and the fix depends on no file outside this repo — which is the exact property the reorder stop-gap lacks. ### Measured `ZDOTDIR` is honoured by a login zsh started from a clean parent: ``` $ env -i HOME=$HOME PATH=/usr/bin:/bin ZDOTDIR=$S /bin/zsh -l -c '...' MARK .zshenv ZDOTDIR=/…/zdot MARK .zprofile ZDOTDIR=/…/zdot MARK .zlogin ZDOTDIR=/…/zdot ``` Then the full design: four generated files in `$ZDOTDIR`, each sourcing the operator's real `$HOME/.z*` counterpart, with the scrub appended to `.zlogin`. Run as an **interactive login shell from a clean parent** (`env -i HOME=… ZDOTDIR=… /bin/zsh -l -i`), so nothing is inherited from my session: ``` fleet: allowed 3 of 28 credential-shaped variables SURVIVES AI_GATEWAY_TOKEN (48) CONTEXT7_TOKEN (64) WORKER_GITEA_TOKEN (40) BLOCKED AWS_ACCESS_KEY_ID · AWS_SECRET_ACCESS_KEY · JENKINS_MCP_AUTH · CONFLUENCE_API_TOKEN · GITLAB_PERSONAL_ACCESS_TOKEN · GITEA_ACCESS_TOKEN · TELEGRAM_BOT_TOKEN · N8N_WEBHOOK_TOKEN · N8N_ENCRYPTION_KEY · N8N_OWNER_PASSWORD · HASS_TOKEN · CF_API_TOKEN · CF_USER_TOKEN · TS_AUTHKEY · TS_API_KEY · BESZEL_ADMIN_PASSWORD · BESZEL_KEY · BESZEL_UNIVERSAL_TOKEN · GRAFANA_ADMIN_PASSWORD · METRICS_PUSH_TOKEN · BRAIN_MCP_TOKEN · MEMORY_MCP_TOKEN · LTMS_API_KEY · HW_PASSWORD · … ``` That is the AWS `AdministratorAccess` key blocked by the daemon's own file, with the operator's `.ltms` reordering removed from the equation entirely. I also checked that no operator rc file sets `ZDOTDIR` (`grep` across `.zshenv .zprofile .zshrc .zlogin .zlogout` → no match), so the daemon's value is not overwritten. ### The one thing this measurement exposes about the pattern list `CONFLUENCE_USERNAME` was **not** blocked. It matches none of `*TOKEN* *KEY* *SECRET* *PASSWORD* *AUTH* *CREDENTIAL*`, because a username is the other half of a credential and is not shaped like one. That is this ticket's own defect in miniature: a pattern list is an enumeration, and it can only catch what somebody's pattern happened to describe. So the policy should be a **real** allow-list: blank every variable that is neither on the derived allow-list nor on a small infrastructure passthrough set (`PATH`, `HOME`, `SHELL`, `TERM`, `LANG`, `TMPDIR`, `USER`, `PWD`, `SSH_AUTH_SOCK`² …). Then a credential nobody has thought of is blocked because it is new, not because it matched. Credential-shaped patterns stay only as a *warning* signal in the log, never as the control. ² `SSH_AUTH_SOCK` is #110 (CB-607) — it is a handle to the operator's agent, not a value in any secret file, which is why three credential tickets missed it. Decide it explicitly here rather than letting it pass through by silence. ### Revised acceptance criteria 1. `memberCredentials:` gains `policy: allow-list`. `known:` becomes reporting-only. 2. The allow-list is **derived** from what the launcher itself injects — `gitTokenEnv`, `gitHostEnv`, `tokenEnv`, and every profile's `env:` map — plus the infrastructure passthrough set. Not hand-typed: a hand-typed list breaks every spawn the first time a profile adds a variable. 3. The daemon **generates the `$ZDOTDIR` directory per spawn** and passes `ZDOTDIR` in `tab.create`'s env map. No operator-owned file is edited, and nothing needs keeping in sync. Each generated file sources its `$HOME` counterpart first so `PATH` and the agent binaries still resolve. 4. Non-zsh shells: if the member's shell is not zsh, `ZDOTDIR` does nothing and the member is **unprotected**. Detect it and refuse the spawn, or log a loud WARN — do not fail silently. Say which was chosen and why. 5. A test that starts a real login shell from a clean parent and asserts the surviving set **equals** the allow-list. Equality, not "these names are blocked" — a blocked-name list is the very thing that fails here. 6. Log the denominator: `allowed N of M`, never `blocked N` (#113). 7. One live spawn as proof. The artefact is a shell file handed to another program; a unit test on the generator proves nothing about zsh (#113 instance 4). ### What this does not change The `.ltms` reorder stays as defence in depth. Rotating the AWS key is still first and still the operator's call.
Author
Owner

The control is built, merged, tested — and switched off on the live daemon

Checked in the code on 2026-08-28, not inferred from the ticket.

fleetd.yaml line 198 reads policy: deny-by-default. So HerdrPeerLauncher.applyEnvironmentAllowListPolicy returns null at its first branch, no ZDOTDIR is generated, and applyMemberCredentialPolicy falls back to overlayBlockedCredentials — the CB-596 pre-shell env overlay. That is the exact control this ticket measured failing, because a file sourced later re-exports over it.

The daemon says so itself at every boot. Tonight, after a restart at 22:31:

WARN memberCredentials gap: 5 credential-shaped env var name(s) are on neither known: nor allow:
     — every member pane inherits them UNBLOCKED —
     [AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH, LLM_KEY_N8N, N8N_WEBHOOK_TOKEN]

AWS_ACCESS_KEY_ID is the AdministratorAccess key from the comment above. So the exposure this ticket was raised for is live, and the fix for it has been sitting merged and inert.

This is the defect family the ticket already names, one level up: not "a control that misses names", but a control that is complete and turned off, with nothing that fails when it is. Nothing in the test suite asserts which policy the shipped config selects, and nothing could — the config is gitignored, so no worker and no CI run has ever seen it.

Why the flip is not a one-line config change

MemberEnvAllowList.derive builds the kept set from every profile's gitTokenEnv, gitHostEnv, tokenEnv and env: keys, plus INFRASTRUCTURE_PASSTHROUGH. applyEnvironmentAllowListPolicy then unions in this launch's own env-map keys and, if configured, SSH_AUTH_SOCK.

It never reads memberCredentials.allow:. The live config lists four names there that are consumed by the member's own binaries, not by fleetd, and none of them is derived:

name reader derived today?
GITEA_HOST the PR step no — no profile sets gitHostEnv:
OPENCODE_AUTOMODE_MODEL opencode no
CONTEXT7_TOKEN the context7 MCP server no
CLAUDE_CODE_MESSAGING_TOKEN claude-code no

All four resolve on this host. Flipping the policy today blanks all four. Two of them are secrets, so the documented escape hatch — put the name in a profile's env: map — is not available: that map holds values, and using it would write a secret into config.

The gap is in the design, not in the config. Derivation covers what fleetd injects. It has no way to express what the member's binary needs from the inherited environment. The config already holds exactly the right thing for that — a list of names — and the code ignores it.

What is being done

  1. Union memberCredentials.allow: into the allowed set under policy: allow-list. Derivation stays the base; the operator's list is additive, so it can only widen the set and cannot break another spawn. That is not the hand-mirrored list this ticket rejected — what was rejected was a hand-typed list replacing derivation.
  2. Fix logCredentialGap. Under allow-list its "inherits them UNBLOCKED" sentence is false: a name on neither list is blanked because it is not on the derived set. A control that reports a false exposure teaches operators to skim warnings (#115).
  3. Then flip the live config, add gitHostEnv: GITEA_HOST to each profile that has gitTokenEnv, restart, and prove it with a live spawn: read the pane's own allowed N of M report, confirm the AWS key is blank, and confirm a member can still push and open a PR.

sshAuthSock stays unset, which means SSH_AUTH_SOCK is blocked — that is #110's answer, and it is safe here because members push over the repo-scoped HTTPS token, not the operator's agent.

One acceptance criterion this episode earns

  1. Something must fail when the shipped config does not select the policy. Criteria 1–7 all test the mechanism; every one of them passed while the mechanism was switched off. The daemon already logs the policy at startup — that is not enough, because it is a line in a log nobody diffs. Make it loud, or make it the default, or assert it somewhere that breaks the build. Decide which, and say why.
## The control is built, merged, tested — and switched off on the live daemon Checked in the code on 2026-08-28, not inferred from the ticket. `fleetd.yaml` line 198 reads `policy: deny-by-default`. So `HerdrPeerLauncher.applyEnvironmentAllowListPolicy` returns `null` at its first branch, no `ZDOTDIR` is generated, and `applyMemberCredentialPolicy` falls back to `overlayBlockedCredentials` — the CB-596 pre-shell env overlay. That is the exact control this ticket measured failing, because a file sourced later re-exports over it. The daemon says so itself at every boot. Tonight, after a restart at 22:31: ``` WARN memberCredentials gap: 5 credential-shaped env var name(s) are on neither known: nor allow: — every member pane inherits them UNBLOCKED — [AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, JENKINS_MCP_AUTH, LLM_KEY_N8N, N8N_WEBHOOK_TOKEN] ``` `AWS_ACCESS_KEY_ID` is the `AdministratorAccess` key from the comment above. So the exposure this ticket was raised for is live, and the fix for it has been sitting merged and inert. This is the defect family the ticket already names, one level up: not "a control that misses names", but **a control that is complete and turned off, with nothing that fails when it is**. Nothing in the test suite asserts which policy the shipped config selects, and nothing could — the config is gitignored, so no worker and no CI run has ever seen it. ## Why the flip is not a one-line config change `MemberEnvAllowList.derive` builds the kept set from every profile's `gitTokenEnv`, `gitHostEnv`, `tokenEnv` and `env:` keys, plus `INFRASTRUCTURE_PASSTHROUGH`. `applyEnvironmentAllowListPolicy` then unions in this launch's own env-map keys and, if configured, `SSH_AUTH_SOCK`. It never reads `memberCredentials.allow:`. The live config lists four names there that are consumed by the **member's own binaries**, not by fleetd, and none of them is derived: | name | reader | derived today? | |---|---|---| | `GITEA_HOST` | the PR step | no — no profile sets `gitHostEnv:` | | `OPENCODE_AUTOMODE_MODEL` | opencode | no | | `CONTEXT7_TOKEN` | the context7 MCP server | no | | `CLAUDE_CODE_MESSAGING_TOKEN` | claude-code | no | All four resolve on this host. Flipping the policy today blanks all four. Two of them are secrets, so the documented escape hatch — put the name in a profile's `env:` map — is not available: that map holds **values**, and using it would write a secret into config. The gap is in the design, not in the config. Derivation covers *what fleetd injects*. It has no way to express *what the member's binary needs from the inherited environment*. The config already holds exactly the right thing for that — a list of **names** — and the code ignores it. ## What is being done 1. Union `memberCredentials.allow:` into the allowed set under `policy: allow-list`. Derivation stays the base; the operator's list is additive, so it can only widen the set and cannot break another spawn. That is not the hand-mirrored list this ticket rejected — what was rejected was a hand-typed list *replacing* derivation. 2. Fix `logCredentialGap`. Under `allow-list` its "inherits them UNBLOCKED" sentence is false: a name on neither list is blanked *because* it is not on the derived set. A control that reports a false exposure teaches operators to skim warnings (#115). 3. Then flip the live config, add `gitHostEnv: GITEA_HOST` to each profile that has `gitTokenEnv`, restart, and prove it with a live spawn: read the pane's own `allowed N of M` report, confirm the AWS key is blank, and confirm a member can still push and open a PR. `sshAuthSock` stays unset, which means `SSH_AUTH_SOCK` is blocked — that is #110's answer, and it is safe here because members push over the repo-scoped HTTPS token, not the operator's agent. ## One acceptance criterion this episode earns 8. **Something must fail when the shipped config does not select the policy.** Criteria 1–7 all test the mechanism; every one of them passed while the mechanism was switched off. The daemon already logs the policy at startup — that is not enough, because it is a line in a log nobody diffs. Make it loud, or make it the default, or assert it somewhere that breaks the build. Decide which, and say why.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#144