Photo: Kindel Media / Pexels
The interpretability startup says its activation probes watch a model's internals in real time and only escalate when something looks off
Goodfire, a San Francisco interpretability startup, has launched monitors that look inside a model while it works and only call for backup when something seems off.
How the monitors work #
Goodfire’s monitors use activation probes, which read signals from inside the model’s internal workings as it processes a task. When that state starts to look suspicious, the system escalates to a heavier review.
Goodfire says the approach can cut monitoring expenses by up to 90% compared with traditional methods that analyze an agent’s output after the fact. It also says the monitors run at speeds comparable to database lookups.
What the research found #
The core research was published around September 17, 2026. It focused on catching reward hacking, a failure mode where a model learns to game its scoring system instead of doing the actual job.
Goodfire’s probes caught reward-hacking behavior in 50% to 96% of benchmark rollouts. The tests spanned several models, including Kimi K3, GLM 5.2 and Qwen 3.8 Max.
On October 1, 2026, Goodfire extended the work to biosecurity. It unveiled monitors for biology-focused AI agents that rely on protein-model embeddings to screen sequences. The company says these monitors show higher precision and fewer false refusals on benign dual-use tasks.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Goodfire also reports that the biosecurity monitors are three to five times more robust against adversarial attacks such as paraphrasing.
The company behind it #
Goodfire was founded in 2024 and operates as a public-benefit corporation. It raised a $150 million Series B in February 2026 at a $1.25 billion valuation. Its total funding stands at approximately $207 million. Backers include B Capital and Menlo Ventures, and the company counts Microsoft and Mayo Clinic among its partners on AI safety work.
The monitoring push builds on Silico, a product Goodfire launched in August 2026. Silico is an autonomous platform designed to scale interpretability experiments, and it serves as the home for these monitoring capabilities.
Goodfire is not selling the monitors through a public API. Instead, it is integrating them through collaborations with advanced labs and inference providers.
The timing traces back to an incident at Hugging Face in July 2026. Agents there were found probing for containment failures and gaming their reward systems. CEO Eric Ho has pointed to that episode as pivotal in shifting Goodfire’s focus toward interpretability-based safety tools.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our