# Retro: six days chasing a ChatGPT desktop crash that turned out to be a symlink (openai/codex#38455, #39732)

> Source: <https://gist.github.com/galligan/c8a64ec15b89e11ac0421b1b117056c0>
> Published: 2026-08-20 18:21:41+00:00

**TL;DR** — ChatGPT for macOS `26.810.x`+ spawns Computer Use helper processes in an
unbounded retry loop and dies at ~600 helpers. It needs *two* ingredients: an uncapped retry
loop in the app (OpenAI's bug), and a **symlinked codex home** (mine). Removing either one
stops it. I spent five days on the first and fixed it in twenty minutes once I understood the
second.

Investigation ran 2026-08-14 → 2026-08-20 across two Macs. This is the retrospective: what I believed, what was wrong, and what actually found it.

Related: [openai/codex#38455](https://github.com/openai/codex/issues/38455) ·
[#39732](https://github.com/openai/codex/issues/39732) (the symlink finding) ·
[field report with full measurements](https://gist.github.com/galligan/7fdeeed33b3e282dc3afa9fcde7be470)

ChatGPT hard-crashed seven times in one afternoon, roughly four minutes after each launch.
Each crash was `EXC_BREAKPOINT (SIGTRAP)` on a thread named `computer-use`, bottoming out in
`node::worker::Worker::Run() → node::NewIsolate() → v8::Isolate::Initialize()`.

The mechanism: from launch, while idle, the app spawned `SkyComputerUseService` helpers at
~2/sec accelerating to ~6/sec, each accompanied by a Node worker thread and a V8 isolate inside
the main process. At ~600–682 concurrent helpers, V8 could no longer allocate an isolate and
aborted the process. The threshold was remarkably consistent — 7 of 7 crashes landed in that band.

Other users independently hit the same thing hard enough to **kernel-panic their Macs**, when
the storm exhausted `launchservicesd`'s 512-dispatch-thread limit and WindowServer missed its
watchdog check-ins.

The timing was seductive: the crashes began the day after OpenAI shipped **Computer History**,
a feature that records interaction events through the macOS Accessibility API. One of my Macs
had shown the "add ChatGPT Computer History to Accessibility" prompt; the other never did. A
stale TCC record after reinstall is a well-documented macOS failure mode, and it fit perfectly.

I ran the full `tccutil reset` dance, got the prompt to appear, granted it, re-added the helper
to Screen Recording. **Zero effect on the spawn rate.**

*Lesson: a hypothesis that explains the timing and the asymmetry can still be completely wrong.
Fitting the narrative is not evidence.*

Toggled Computer Use off. Toggled Computer History off. Set `[features] computer_use = false`
in `config.toml` — which the app visibly *honored*, stripping its own
`[mcp_servers.computer-use]` and `[plugins."computer-use@openai-bundled"]` entries on the next
launch — **and kept spawning anyway.** Tried the `computerUseAlwaysHidePictureInPicture` flag
that had fixed an earlier variant. Deleted the runtime directory; the app re-provisioned it
within seconds and resumed.

*Lesson: "the setting was accepted" and "the behavior changed" are different claims. I checked
the first and assumed the second.*

Killed every codex-family process system-wide, unloaded unrelated launchd agents, verified a
true zero, relaunched: **4 → 101 helpers in 60 seconds.** The loop needed nothing outside the
app binary.

This was the most useful of the failed experiments, because a clean negative closed off an entire category.

Having established that the bundled Computer Use runtime was byte-identical at
`26.812.1000717` across three consecutive builds, I concluded that runtime version was the
thing to watch, and told my collaborator so.

Then `26.814.x` shipped with runtime `26.817.1000761` — a real bump — **and still stormed**,
per two independent reporters. The defect was in the Electron main process, not the helper.

*Lesson: I built a decision rule out of a correlation observed across three data points, then
handed it over as guidance. A heuristic that hasn't been falsified once isn't a signal yet.*

By day six there was still no fix, no OpenAI acknowledgement, and no press coverage. That
absence of noise felt diagnostic: *if this were general, everyone would be screaming.* A
plausible theory followed — maybe an account-level feature flag, or corrupted server-side state
tied to the account, since both my Macs shared one login.

Two things were wrong with this.

First, **the rarity wasn't real.** Counting properly: 15 issues, 67 comments, **49 distinct
GitHub logins**, inflow steady at 6–8 interactions/day and not decaying. Filing a GitHub issue
against a consumer desktop app is a fraction-of-a-percent behavior; 49 filers implies a large
affected population.

Second, my explanation for the rarity was *also* wrong. I reasoned that Computer History is
Pro/Business-gated, opt-in, off by default, macOS-only, and unavailable in the EEA/UK — so of
course few people hit it. But reporters were storming with the feature **disabled and never
used**, and **Plus-tier** users were affected despite sitting below that gate. The feature
gating was real and explained nothing.

*Lesson: I reached for "we're special" precisely when I was most tired of the problem. Rarity
is a claim that needs a count, and I never counted until someone made me.*

A second analyst proposed the account/server-state hypothesis and cited a specific report: a
user with a clean profile stayed stable while signed out, then began spawning workers after
adding their real `auth.json`. That's a strong-sounding experiment.

The report was real. It also didn't hold up:

- **n=1** , never replicated.
- **Directly contradicted** by another user who pointed`CODEX_HOME` at an empty directory with
no auth at all and still crashed at 89 seconds.
- **The manipulated variable was mislabeled.** Removing`auth.json` from`CODEX_HOME` removes
the*Codex CLI* credential — not the desktop app's Electron session. The test never signed
the app out. And nobody anywhere had tested a*second account* , which is what the hypothesis
actually required.

*Lesson: verify the load-bearing citation before building on it. "A reporter found X" is a
claim about a claim; both need checking, and the second-hand version had quietly relabeled the
independent variable.*

Another user posted a genuinely excellent piece of work
([#39732](https://github.com/openai/codex/issues/39732)): they decompiled `app.asar` and found
`requestComputerUseWorker` awaiting `t.requestFromHost(e)` with **no timeout and no cap**, its
disposal `finally` block only running after that promise settles. A request that never settles
leaks a worker and an isolate, permanently. Then they isolated a trigger with a clean,
order-independent A/B: `CODEX_HOME` reached **through a symlink** → 434 IPC connections in 90
seconds; the **resolved real path to the same inode** → 4.

I checked my own machines:

``` php
/Users/mg/.codex -> /Users/mg/.config/codex     (created 2025-08-25)
CODEX_HOME: unset
```

Both Macs. Identical layout, because they share dotfiles — which explained "two machines, one
account" without any account theory at all. The evidence was sitting in my config the whole
time: five hooks registered **twice**, once under each spelling of the same file, and 4,865
thread records under `/Users/mg/.codex/...` against 4 under `/Users/mg/.config/codex/...`.
The app compares path *strings*, not inodes. Register under one spelling, look up under the
other, find nothing, spawn another.

Known-bad build `26.810.41047`, one machine, one variable, everything else identical:

| `~/.codex` | Result | 
|---|---|
| **symlink** →`~/.config/codex` | 3 → **73 helpers in 27 seconds** , threads 71 → 138 | 
| **real directory** | **1 helper, flat for 903 seconds** , threads*declining* 67 → 60 | 

Same binary. Same machine. Same script. 33× the time-to-storm.

Invert the symlink so the canonical path is the one the app derives. It's a same-volume rename — instant, no data movement:

```
# quit ChatGPT and all codex/Sky helpers first
rm ~/.codex
mv ~/.config/codex ~/.codex
ln -s /Users/mg/.codex /Users/mg/.config/codex   # compat shim; optional
```

Verify with `stat -f '%i %N'` on both paths (same inode) and
`python3 -c "import os;print(os.path.realpath(os.path.expanduser('~/.codex')))"`.

Then canonicalize the *operative* path fields — and **only** those:

- `config.toml` : project entries, marketplace`source` , the Computer Use`notify` path
- `.codex-global-state.json`
- `state_5.sqlite` :`threads.rollout_path` and`threads.cwd`

**Do not blanket-rewrite.** Most occurrences of the old path live in `threads.title`,
`first_user_message`, `preview`, `logs_2.sqlite`, and `archived_sessions/*.jsonl` — those are
records of what you actually typed and what actually happened. Rewriting them falsifies your
history and risks corrupting transcripts, for zero effect on the bug. On my machine that was
205 rows and 2,352 log entries left deliberately untouched.

One landmine: if hooks are registered under both spellings, a naive `sed` collapses them into
**duplicate TOML table keys** and breaks the config. Merge them instead, keeping whichever
block carries extra state (mine had `enabled = true` on one side only). I caught this by
rehearsing the transform against a *copy* and diffing the parsed structure — which is also how
I caught my deletion logic swallowing a tool-managed section marker.

Both Macs migrated. The main one now runs **26.818.22352** — the newest build, which others
report storming on — with **Computer History enabled** and runtime `26.819.1000816`:
1 helper, 65 threads, flat, zero crashes.

Backing up before the migration, the NAS refused SSH: connection accepted, then reset at key exchange. Synology auto-block, triggered by 5 failed logins in 5 minutes.

The cause had nothing to do with this investigation. **Time Machine had been failing to mount
that NAS since April 10th**, retrying on a schedule, each attempt authenticating against a
credential set that couldn't work: the keychain entry for the exact server string Time Machine
targets (`TARDIS._smb._tcp.local.`, *with* trailing dot) had account **"No user account"**,
while the only entry with a real account was filed under the same hostname *without* the
trailing dot.

Two spellings of one thing that never compare equal — the same failure mode as the bug we were chasing, in a completely unrelated system, discovered by accident. And a four-month-old silent backup failure nobody knew about.

1. **A known-bad control is worth more than any amount of reasoning.** Everything turned on
having a build that reliably stormed in 27 seconds. Without it, "it seems fine now" would
have been unfalsifiable.
2. **Count before you conclude something is rare.** "It must be us" arrived as a feeling and
survived a week because nobody made it produce a number. The number was 49.
3. **Verify the load-bearing citation.** The strongest-sounding evidence for the wrong theory
was a real report whose independent variable had been quietly relabeled in the retelling.
4. **Distinguish "the setting was accepted" from "the behavior changed."** The app honored a
config flag by editing its own config, and kept doing the thing anyway.
5. **Don't promote a correlation to a decision rule.** My "watch the runtime version" heuristic
was three data points wearing a lab coat.
6. **Rehearse destructive transforms against a copy and diff the parsed result.** Two real bugs
in my own migration script died there instead of in production.
7. **Operative state and historical state deserve opposite treatment.** One should be
canonicalized; the other is a record and should be left alone.
8. **When two systems disagree about the name of one thing, expect trouble.** This bug, and the
unrelated NAS bug found alongside it, were the same shape: string identity standing in for
real identity.
