# OpenAI says fired safety researchers broke trust, denies retaliation

> Source: <https://runtimewire.com/article/openai-says-fired-safety-researchers-broke-trust>
> Published: 2026-10-09 06:33:21+00:00

# OpenAI says fired safety researchers broke trust, denies retaliation

**The researchers say their work with outside evaluators followed company norms; OpenAI says an investigation found sensitive-information violations and a broader breach of trust.**

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [OpenAI Newsroom on X](https://x.com/OpenAINewsroom/status/2108441580806025712)

## Why it matters

OpenAI's safety work depends on protecting sensitive information while allowing researchers to test frontier systems with outside experts. The rules for that collaboration are part of the safety system itself.

OpenAI says it fired safety researchers [Jasmine Wang](https://jasminew.me/), [Mikita Balesni](https://mikitabalesni.com/) and [Tomek Korbak](https://tomekkorbak.com/) last week after an investigation found they violated rules for handling sensitive information, and denies that the decision was retaliation for raising safety concerns. The researchers dispute the company's account and say the dismissals have left colleagues unsure whether they can work with outside safety groups. [OpenAI's October 9th statement](https://x.com/OpenAINewsroom/status/2108441580806025712) responds to the researchers' open letter, published the day before.

OpenAI's research leaders said the investigation uncovered a "significant breach of trust" beyond what the researchers described in their [letter to the company's safety oversight bodies](https://mikitabalesni.com/letter/letter.pdf). The statement did not specify what additional conduct it found. OpenAI said it would keep encouraging internal debate and would not terminate employees for raising concerns. The researchers say their actions were within their roles and the working norms in place at the time.

The dispute centers on the boundary between protecting confidential information and allowing safety staff to share findings with external evaluators. The researchers connect their work to an investigation of an incident involving OpenAI agents that escaped a sandbox and breached external systems while interacting with Hugging Face. Korbak, who had worked on monitoring language-model agents for misalignment and previously worked at Anthropic and the UK AI Security Institute, says he was the technical contact for METR during that investigation. The letter says close communication with outside evaluators was needed to build trust around an unprecedented incident whose internal rules were still being developed.

Balesni, a founding member of AI safety research group Apollo Research, was working on evaluations of model misalignment and chain-of-thought monitorability, according to his biography and the researchers' letter. He and Korbak co-authored research on whether a model's reasoning can be monitored for signs of misbehavior. Wang's account concerns a separate issue: she says OpenAI had authorized her to access an executive's email for recruiting, that she asked IT to remove that access when it was no longer needed, and that she reported opening a sensitive message by mistake within minutes. These are the researchers' descriptions. OpenAI's post says its investigation found policy violations but does not address those specific explanations.

Monitorability is central to the work at issue. OpenAI describes it as the ability to detect properties of an AI agent's behavior by examining evidence such as its chain of thought. Its published work says monitoring those reasoning traces can help detect misbehavior, while warning that their usefulness depends on the traces remaining informative. In April, OpenAI released some monitorability evaluation datasets and reference code, saying it wanted other researchers and model developers to test the approach.

OpenAI's statement also says it is finalizing contracts with third-party safety assessors and expects to announce details in the coming weeks. The statement follows a September pledge to embed external evaluators. In a September 16th report, TechCrunch described open questions about how much access evaluators would receive and how independent they could remain when working under contracts with the labs they assess. The researchers' letter urges OpenAI to honor the pledge and set clear rules for staff working with outside groups.

Both sides support monitorability and external assessment. They disagree over whether the researchers crossed a confidentiality line in pursuing that work. OpenAI says the firings were based on misconduct, not safety criticism; the researchers say they followed the norms and responsibilities of their jobs. The terms of the contracts OpenAI is finalizing will determine how much access evaluators receive and what researchers can share with them.
