cd /news/artificial-intelligence/training-ai-scientists-to-replicate-… · home topics artificial-intelligence article
[ARTICLE · art-97128] src=inherentlabs.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Training AI Scientists to Replicate Research

Researchers introduced Faraday, a 27B-parameter AI Scientist agent that outperforms Claude Opus 4.8 and GPT-5.5 on replicating research figures, trained via long-horizon RL on the Replica task suite of 310 tasks from 100 ML and AI for science papers. The team validated an LLM judge with a human study and used per-task rubrics, multi-sample aggregation, and turn-level credit assignment to stabilize training, aiming to advance AI toward scientific innovation.

read4 min views1 publishedAug 14, 2026
Training AI Scientists to Replicate Research
Image: source

thAugust 2026

Read the paper↗ In 1821, Richard Phillips asked his friend Michael Faraday to write a review of the emerging field of electromagnetism. Faraday was a strange choice for the task. The British scientist had little formal education and limited experience in the field. But he was a well-known experimentalist.

To write his review, Faraday decided to replicate past results manually. In the course of his experiments from the candlelit basement of the Royal Institution, Faraday found he could make a wire carrying current move around a magnet. “Very satisfactory”, he wrote in his journal. Faraday had invented the electric motor.

Inspired by Michael, we introduce Faraday, a 27B-parameter “AI Scientist” agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research. Trained via long-horizon RL with coding agents as a tool, Faraday learns the skills of a rigorous scientist, a step towards AI Scientists capable of innovation across domains.

From replication to innovation #

To train Faraday, we introduce Replica, a scalable space of RL tasks. Each task requires an agent to replicate a figure from a research paper with a limited time and compute budget, and without access to the original plot. The initial suite comprises 310 tasks from 100 ML and AI for science papers, in domains as diverse as natural language processing, materials science and weather forecasting.

We run recent agents from the leading labs on Replica and find that they do not saturate the task space. For Claude Opus 4.8 and GPT-5.5 baselines, we run the model in the Claude Code and Codex harnesses respectively, with thinking effort set to extra high. Faraday produces more faithful replications for every category of paper in the task suite, and struggles less with recent research, effectively applying its scientific skills to work unseen by the base model at pre-training.

Replicating a figure may not seem like an especially innovative task. But looking at the process, rather than the output, replication becomes a stepping stone towards innovation. Research papers describe what the authors found that worked, not the negative results that got them there. To succeed, agents must recover the “99% perspiration” that does not appear on the page. This requires the hypothesis-driven exploration characteristic of open-ended research.

Replication forms the basis for a curriculum of underspecified tasks. Features beyond single plots can be removed from papers given to agents to replicate. Resource constraints can be further tightened or relaxed. Papers can even be imagined, leading the same Faraday model to innovate without knowing it.

RL for non-verifiable domains #

Perfectly reproducing a plot is not the same as a successful replication, which also requires strong experimental design, good scientific practice, faithfulness to the claims of the original paper, and effective use of available resources. Ultimately, replication requires “research taste”.

We design an LLM judge and run a human study to validate that it captures the research taste of experts. But training on LLM judges remains challenging, because the stochasticity of the underlying model creates a noisy reward signal. We use per-task rubrics to solve this problem, achieving greater consistency and lower noise than an LLM baseline.

To address the instability typical of long-horizon RL training, we make two further train-time modifications to our judge: multi-sample aggregation and turn-level credit assignment. Trained using this recipe, Faraday learns to become a more rigorous scientist.

Scalable scientific oversight #

Faraday employs GPT-5.5 Codex as a tool, much like human scientists use coding agents. Remarkably, Faraday directs the work of a model several orders of magnitude larger in a way that improves replication performance.

Moreover, Faraday can generalise to directing a more capable agent at test-time, adapting to GPT-5.5 Codex after training with GPT-5.4-mini. As frontier coding agents continue to advance, we expect that the value of scientific judgement will only increase.

Faraday builds on previous discoveries to find new insights at test-time, similarly to existing AI Scientist agents. But unlike these agents, Faraday requires no hand-coded evolutionary harness, and has no test-time reward. In other words, Faraday learns to value discoveries intrinsically.

Future directions #

Faraday represents the first step in a new paradigm that combines a layer of scientific intuition with the advancing capabilities of coding agents. We believe that better AI Scientists, powered by novel infrastructure and new forms of human-machine teaming, can benefit all of society.

At Inherent, Faraday’s improvements compound through the entire company, enabling us to discover new knowledge while firmly keeping humans in the loop. With an eye to safety, we are investigating how our methods might advance scalable oversight and ameliorate reward hacking.

We are building a new kind of AI and a new kind of research institution, fit for the age of AI-driven scientific discovery. To join us on our mission, apply here.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @faraday 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/training-ai-scientis…] indexed:0 read:4min 2026-08-14 ·