Seeding ~/.claude.json can still lose an external writer's change #247
Closed
opened 2026-09-03 06:47:57 +02:00 by ltms
·
3 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#247
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up to #149, which is merged (
2e5b63f). #149 made the workspace-trust seed safe inside fleetd: the read-modify-write is atomic (ATOMIC_MOVE) and guarded by a process-wide lock, so two concurrent spawns can no longer lose each other's entry.What is still broken. The lock is in-process.
~/.claude.jsonis also written by the operator's own running Claude Code, which fleetd cannot see or coordinate with. So this can still happen:The move is atomic, so the file is never torn. But
v3was built fromv1, so their change is gone. The window is small (one read plus a small edit), and it only opens when a spawn happens at the same moment Claude Code saves. It is real, not theoretical: the file is written on ordinary events like adding a project or finishing an OAuth refresh.Why it was not fixed in #149. Fixing it needs a cross-process mechanism, and none of the options is obviously right. I chose to merge the in-process fix and file this rather than hold a third review round.
Options, roughly in order of cost:
~/.claude.json.lock,O_CREAT|O_EXCLorflock). This actually closes the window, but only if the other writer takes the same lock — and Claude Code does not. So on its own it buys nothing against the real writer. It is only worth doing if it is paired with (1).Suggested first step: check whether option 3 is possible before building 1 or 2. If it is not, do 1, and say plainly in the code comment that the window is narrowed and not closed.
Acceptance criteria:
~/.claude.json, or the seed detects an external write and retries.Measured live on 2026-09-03 after deploying #149. The external writer is worse than I described when I filed this, and it changes which fix is right.
I spawned a real
claude-codemember into a fresh worktree and then read<configDir>/.claude.json:ClaudeCodeLauncher.java:561-562writes both keys, and our entry was created by the running jar minutes before I looked. SohasCompletedProjectOnboardingwas written and is gone — not only from our entry, but from every entry in the file.What that means
I filed this as a race: a save landing in the window between our read and our
ATOMIC_MOVE, losing whatever changed in between. That framing is too kind. Claude Code normalises the whole file on save and drops keys it does not expect in that position. It is not losing a race with us; it is deliberately rewriting what we wrote, every time, with no window involved.This rules out option 1
Compare-and-swap on content was the cheap option in this ticket, and it is now clearly the wrong one. CAS defends against a concurrent write landing between a read and a write. It does nothing when the other writer's steady-state behaviour is to remove your key. A retry loop would re-add
hasCompletedProjectOnboardingand the next save would strip it again — an endless, invisible fight over a key that does nothing.Option 2 (an advisory lock) was already weak, because Claude Code does not take our lock. This measurement does not rescue it.
What is left, and it is simpler than the original three options
Stop writing the second key. We now know two things about
hasCompletedProjectOnboarding: it is not retained, and it is not needed — the member reachedidlewith onlyhasTrustDialogAcceptedpresent (see #149). Writing it accomplishes nothing and is the only part of our write that provably conflicts with the other writer.That leaves
hasTrustDialogAccepted, which does survive — all 28 entries carry it, so Claude Code preserves that key rather than stripping it. The narrow race on that one key is still theoretically real, but its blast radius is one boolean that both writers agree should betrue, which is the benign case.So the revised suggestion, replacing the three options above:
hasCompletedProjectOnboardingwrite. One line.Acceptance criteria, revised:
hasTrustDialogAccepted.idle. That is the only thing the write exists to achieve.The original option 3 — find a way not to write the file at all — is still the ideal, and still unexplored. But it is now much less urgent, because what remains is a single boolean the other writer already agrees with.
This ticket names the wrong file, and that made the race look far less likely than it is. Measured on the live host today.
Option 3 is already built, and it does not solve the problem
ClaudeCodeLauncheralready exportsCLAUDE_CONFIG_DIRto the member fromcfg.configDir()(line 284) and already seeds<configDir>/.claude.jsonrather than~/.claude.jsonwhen that is set (lines 531-533). So "point fleetd at a file it owns" needs no new code — only config.And the live config already does it. Every claude-code profile sets
configDir:(
sol,terra,xf,gxare opencode and never seed a trust dialog.)So in production fleetd has never written
~/.claude.json. That part of the premise is wrong.But the file it does write is not one fleetd owns either
CLAUDE_CONFIG_DIRfor the operator's own lead session is/Users/dai.ha/.ccs/instances/ltms— the same directory thesonnetandopusprofiles point at. Its.claude.jsonis 125528 bytes with 35 projects, 8 of them fleetd worktree entries.Its mtime moved to
16:03:21today, minutes after a spawn, while the operator's Claude Code was running and writing the same file.So
configDirdid not move the write onto a file nobody else touches. It moved it from one live file to another live file — the lead's own. The race described in this ticket is real and live; it is just not on~/.claude.json. Naming that file made it look like a rare collision with an unrelated process. It is actually a collision with the session that is doing the spawning.What this probably explains
~/.claude.jsonwas corrupted today: it shrank to 20208 bytes and lostoauthAccount— a lost update, not a torn write, which is this ticket's exact signature.The path there is #258: two test fixtures built a worktree-shaped
@TempDirwithconfigDirnull, so the test suite — not production — wrote the real~/.claude.jsonon every run. 116 dead JUnit entries had accumulated there. A test run reading v1 while the operator's Claude Code wrote v2 produces precisely the observed damage.I cannot prove this. The corrupted copy is no longer on disk, so the evidence is gone. But it is the only mechanism I found that produces that signature, and both halves of it are independently confirmed: #258 (fixed,
ac790e4) put the test path on the real file, and this ticket is why any writer of that file can lose a concurrent change.Revised recommendation
Option 3 as written — "find another way to mark a worktree as trusted" — is not available.
claude --helpoffers no flag to accept workspace trust. The dialog is skipped only in non-interactive mode (-p, or a non-TTY stdout), and a member runs interactively in a pane, so that does not apply. The per-project trust flag lives only inprojects.<cwd>.hasTrustDialogAcceptedin a global-per-config-dir file.Pointing
configDirat a directory fleetd genuinely owns is not a fix either: the member is given that same directory as its ownCLAUDE_CONFIG_DIR, and it is set to the operator's instance dir precisely so the member inherits that instance's configuration.So: do option 1 — re-read immediately before the
ATOMIC_MOVE, and start over if the bytes changed since the read. And say plainly at the call site that this narrows the window and does not close it, because a write landing between the re-read and the move is still lost.Two things the acceptance criteria should add, given the above:
~/.claude.json" will assume it is rare. The truthful sentence is that fleetd writes the lead's own config file, on every claude-code spawn.configDirunset is the only path to~/.claude.json, and it is the default. Nothing warns. #258 shows what that costs when something reaches it by accident — 116 writes to an operator's file with a green build every time.Fixed. PR #261 is merged to
mainas34480ce. Full build: 1261 tests, 0 failures.What shipped
seedTrustDialognow does a compare-and-swap with a bounded retry, inside the existingTRUST_JSON_LOCK:ATOMIC_MOVE, re-read the bytes and compare. If they changed, someone else wrote in between — throw the attempt away and rebuild from the fresh bytes.Step 4 is the important one. An unseeded member shows the trust dialog and fails to reach an injectable state. That is visible, logged and recoverable — you retry the spawn. Writing our stale copy over the operator's live config is silent and not recoverable. So the code fails toward the recoverable outcome.
There is also a new WARN every time the seed is about to target the default
~/.claude.json(configDirunset). That path is the only one that touches the operator's home file, and it is the default, so a profile that simply forgot to setconfigDirpreviously got no signal at all.The ticket named the wrong file, and I should correct that
I filed this saying the file at risk was
~/.claude.json. That is not where the live race is.The real race is against
<configDir>/.claude.json. On this host that resolves to/Users/dai.ha/.ccs/instances/ltms/.claude.json— the lead's own live config file, being read and written by the very session doing the spawning. I confirmed it from the config:local,local-direct/Users/dai.ha/.ccs/instances/gx10opus,sonnet/Users/dai.ha/.ccs/instances/ltmsgx,sol,terra,xfSo this was never an exotic collision. It was a lost update on every ordinary spawn of
opusorsonnet. Everyclaude-codeprofile has aconfigDir; the four without one are allopencode, which never calls this method at all.What this does NOT fix — stated plainly
The window is narrowed, not closed. A write landing between the final re-read and the
ATOMIC_MOVEitself is still lost. There is no OS-level compare-and-swap on a plain file, only this cooperative narrowing.That limit is written into the method's javadoc, not just here, so the next person to read the code does not assume the race is gone and stop looking.
Proof the guard is load-bearing
The worker removed the CAS check (
if (false && !Arrays.equals(before, atMove)), i.e. move unconditionally, as before the fix) and both new race tests went red. The failure message is worth quoting because it says what was actually lost:The file was then restored and verified byte-identical before committing.
Was this the cause of the corruption on 2026-09-03?
I still cannot say it was, and I am not going to claim it. No
~/.claude.json.CORRUPTED.*file exists on disk any more, so there is nothing left to examine. What I can say is that the mechanism was real, it was reachable on everyopus/sonnetspawn, and it is now bounded.Closing.