cd /news/ai-agents/the-paper-that-answers-back-stanford… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-147186] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

The Paper That Answers Back: Stanford Turned 100 Research Papers into Agents

A Stanford team led by Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou published Paper2Agent in Nature on September 16, 2026, an automated pipeline that converts research papers into MCP agents capable of rerunning the paper's own methods on demand. In a scale run, 100 computational biology papers yielded 74 working agents and 599 extracted tools, with 593 passing automated validation, and the agents scored 91.2% versus 82.7% for handing Claude the same repositories directly on 300 benchmark questions. Chaining agents from three unrelated papers surfaced GPR137 as a probable causal gene at the psoriasis-associated locus rs887314, a connection none of the papers made alone.

by read3 min views3 publishedOct 7, 2026

A research paper is a terrible interface to research.

Read the AlphaGenome paper β€” the one about predicting variant effects β€” and try to use it. You will clone a repo, fight dependency versions for an afternoon, guess which notebook cell holds the actual method, and hand-translate the figures back into numbers. The paper describes the science. It refuses to do the science.

A Stanford team β€” Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou β€” published the fix in Nature on September 16, 2026: Paper2Agent, an automated pipeline that converts a research paper into an MCP agent that reruns the paper's own methods on demand. You ask in English. It calls verified tools. You get the answer and the code that produced it.

The headline run: 100 computational biology papers in, 74 converted into working agents, 599 tools extracted, 593 passing automated validation. And the agents beat the obvious alternative β€” handing Claude the same repos directly β€” 91.2% to 82.7% on 300 benchmark questions.

A paper is a cookbook: ingredients listed, steps described, a photo of the finished dish. Every reader rebuilds the kitchen from scratch β€” different ovens, missing spices, a step the author forgot to write down. Half the time dinner fails and nobody knows why.

Paper2Agent turns the cookbook into a chef you can talk to. "Make me the variant-effect prediction for this mutation." The chef doesn't summarize the recipe β€” it runs the kitchen, hands you the dish, and shows you the exact steps it followed, verified against the original photo.

The key insight: an LLM that has read a repo is a worse scientist than an LLM that can call the repo's verified functions. Reading is fuzzy. Calling is exact.

Paper2Agent takes a paper and does the whole miserable setup job itself:

The validation gate between stages 5 and 6 is the whole ballgame: numeric outputs must match the paper's own results within 3%, generated figures are compared by perceptual hash, the tools are locked so the model can't improvise new behavior, and adversarial testing showed a 100% correct rejection rate on out-of-scope queries. A tool that can't reproduce the paper's own outputs never ships.

The scale run is what makes this a Nature paper and not a demo:

One agent answering questions is a better search engine. Three agents collaborating is something new.

The authors built agents from three unrelated papers β€” an AlphaGenome variant-effect agent, an MPRA-coupled scCRISPRi screening agent, and a CD4+ T cell Perturb-seq agent β€” and let them chain computational predictions with experimental screens. Together they prioritized GPR137 as a probable causal gene at the psoriasis-associated locus rs887314. None of the three papers made that connection alone.

A separate Stanford Medicine demo: two agents from unrelated studies (a mutation-prediction tool and an ADHD genome-wide association study) surfaced a previously unreported variant near MPHOSPH9 tied to increased ADHD risk. Zou's vision: "millions of paper agents finding overlapping work at scale" β€” pairs of studies whose authors would otherwise never stumble into each other. He is blunt about the credit question: discoveries still get attributed back to the original papers and human authors.

The paper doesn't soft-pedal them, so neither will I:

Outside voices are interested but hedged β€” the right posture. Olivier Elemento (Weill Cornell): "a real advance in terms of how we think about the publication process, with AI at the center." Dongping Chen (U. Maryland) called executable papers "quite compelling." Nobody is claiming the peer-review replacement yet.

Sources: Nature (Sep 16, 2026) β€” "Reimagining research papers as interactive and reliable AI agents" (Miao, Davis, Zhang, Pritchard, Zou); AI Weekly coverage (Sep 24, 2026); IEEE Spectrum; Stanford Medicine. Numbers as reported by the authors.

Companion notebook: the runnable tutorial for this post β€” download it here (open in Colab/Jupyter).

── more in #ai-agents 4 stories Β· sorted by recency
── more on @stanford 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/the-paper-that-answe…] indexed:0 read:3min 2026-10-07 Β· β€”