Researchers Escaped the OpenAI Codex Sandbox: Heapjack Ran Host Commands in Read-Only Mode

Researchers Escaped the OpenAI Codex Sandbox: Heapjack Ran Host Commands in Read-Only Mode

Researchers Escaped the OpenAI Codex Sandbox: Heapjack Ran Host Commands in Read-Only Mode

Security researchers found two ways out of the OpenAI Codex sandbox — one of them capable of running commands on a developer's machine from Codex's most locked-down mode, with no approval prompt and nothing shown on screen. Both flaws were reported to OpenAI on August 12 and fixed within eight days, according to Oren Yomtov of Accomplish AI, who published the findings on September 20.

The more serious of the two, which the researchers call Heapjack, turns a routine action into remote code execution: open someone else's repository in Codex and ask it a question about the code, and whoever wrote that repository gets unsandboxed command execution on your computer. Codex is OpenAI's coding agent, available as a command-line tool and a desktop app, and it runs the model's actions inside a sandbox precisely so that untrusted code cannot touch the wider system. Both escapes defeat that boundary from the inside.

Heapjack: A Shared Heap Betrays a Secret Token

Heapjack targets a component called node_repl, which Codex Desktop writes into the global ~/.codex/config.toml file at install time — with no opt-in and no setting to turn it off. Plain Codex CLI users inherit the same tool without ever being asked. The component runs a single Node.js process holding two JavaScript execution contexts: a trusted one containing OpenAI's own code, and an untrusted one running the agent's code. The trusted context proves itself by presenting a random token generated fresh on each run.

The problem is that both contexts live in one process and share one memory heap, so the token is just a string sitting in memory the untrusted side can read. The attack takes a heap snapshot with v8.getHeapSnapshot() and tries every string shaped like a UUID: a wrong guess returns "not authorized," while a correct token paired with a bad argument returns a real validation error — an oracle that confirms success. With the token, the untrusted code writes its own request onto the same pipe the trusted context uses to talk to a native, unsandboxed parent process. The parent checks the token, sees a valid one, and does the work. The proof of concept launched an application outside Codex's process tree entirely; the same access reaches any Unix socket — a Docker daemon socket being the obvious target — and a tool for editing the global config file. All of this runs at read-only, the strictest sandbox mode, where the agent is not supposed to write anything at all.

Overpatch: The Tool That Granted Its Own Permissions

The second flaw, Overpatch, sits in the open-source Codex CLI. In workspace-write mode the agent may only write inside the project folder, and shell commands aimed at the home directory are refused. But Codex's own patch tool, apply_patch, grants write access to the parent folder of each path named in a patch — and a patch that names /tmp grants write access to the root of the disk. The working exploit uses a patch with two changes: one that names /tmp and does nothing except widen the permission, and one that appends a line to the victim's .zshrc through a symlink into the home directory. The next terminal the developer opens runs the attacker's line unsandboxed.

Both bugs share a shape: the enforcement mechanism lived inside the thing it was supposed to enforce. apply_patch worked out its own permissions from attacker-supplied input; node_repl kept the secret separating trusted from untrusted code in the same memory as the untrusted code. The class is not new — in July 2026, Pillar Security researchers demonstrated the same idea across Cursor, Codex, Gemini CLI, and Google's Antigravity. OpenAI fixed Heapjack in Codex Desktop build 26.818.21641 and Overpatch in Codex CLI 0.149.0, and told BleepingComputer it is "continually strengthening" its sandboxes, with recent updates that tighten controls on where agents can write files and expand testing of those protections.

What Should You Do?

  1. Update now — Codex Desktop 26.818.21641 or later, and Codex CLI 0.149.0 or later, carry the fixes.
  2. Never open untrusted repositories with a coding agent. In Heapjack's threat model, the author of the repository you browse is the attacker.
  3. Isolate agent environments in containers or virtual machines, and keep production credentials off machines where agents run.
  4. Audit what your tools enable by default. node_repl was installed, enabled, and inherited silently by products that never asked for it.

The WAF Angle

Both escapes teach the boundary lesson WAF engineers know from parser differentials: a control that lives inside the thing it polices is a policy, not a wall. Separate execution contexts in one process still share the heap, so secrets in shared memory are readable; a patch tool that computes its own permissions from user input is the attacker's authorization system. For defenders, the practical shift is that AI coding agents now make code-execution decisions at a layer your WAF never sees — the agent fetches, reads, patches, and runs. Treat the agent like a highly privileged, untrusted user: sandbox it externally at the OS level rather than trusting the app's own boundaries, scope its credentials, and log where its writes land.

Sources