# The Paper That Answers Back: Stanford Turned 100 Research Papers into Agents

> Source: <https://dev.to/danielsamfdo/the-paper-that-answers-back-stanford-turned-100-research-papers-into-agents-55go>
> Published: 2026-10-07 21:54:07+00:00

A research paper is a terrible interface to research.

Read the AlphaGenome paper — the one about predicting variant effects — and try to *use* it. You will clone a repo, fight dependency versions for an afternoon, guess which notebook cell holds the actual method, and hand-translate the figures back into numbers. The paper describes the science. It refuses to *do* the science.

A Stanford team — Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou — published the fix in **Nature on September 16, 2026**: **Paper2Agent**, an automated pipeline that converts a research paper into an MCP agent that reruns the paper's own methods on demand. You ask in English. It calls verified tools. You get the answer *and the code that produced it*.

The headline run: **100 computational biology papers in, 74 converted into working agents, 599 tools extracted, 593 passing automated validation.** And the agents beat the obvious alternative — handing Claude the same repos directly — **91.2% to 82.7%** on 300 benchmark questions.

A paper is a cookbook: ingredients listed, steps described, a photo of the finished dish. Every reader rebuilds the kitchen from scratch — different ovens, missing spices, a step the author forgot to write down. Half the time dinner fails and nobody knows why.

Paper2Agent turns the cookbook into a **chef you can talk to**. "Make me the variant-effect prediction for this mutation." The chef doesn't summarize the recipe — it *runs the kitchen*, hands you the dish, and shows you the exact steps it followed, verified against the original photo.

The key insight: an LLM that has *read* a repo is a worse scientist than an LLM that can *call* the repo's verified functions. Reading is fuzzy. Calling is exact.

Paper2Agent takes a paper and does the whole miserable setup job itself:

The validation gate between stages 5 and 6 is the whole ballgame: **numeric outputs must match the paper's own results within 3%**, generated figures are compared by **perceptual hash**, the tools are **locked** so the model can't improvise new behavior, and adversarial testing showed a **100% correct rejection rate** on out-of-scope queries. A tool that can't reproduce the paper's own outputs never ships.

The scale run is what makes this a Nature paper and not a demo:

One agent answering questions is a better search engine. *Three agents collaborating* is something new.

The authors built agents from three unrelated papers — an AlphaGenome variant-effect agent, an MPRA-coupled scCRISPRi screening agent, and a CD4+ T cell Perturb-seq agent — and let them chain computational predictions with experimental screens. Together they prioritized **GPR137 as a probable causal gene at the psoriasis-associated locus rs887314**. None of the three papers made that connection alone.

A separate Stanford Medicine demo: two agents from unrelated studies (a mutation-prediction tool and an ADHD genome-wide association study) surfaced a previously unreported variant near **MPHOSPH9** tied to increased ADHD risk. Zou's vision: *"millions of paper agents finding overlapping work at scale"* — pairs of studies whose authors would otherwise never stumble into each other. He is blunt about the credit question: discoveries still get attributed back to the original papers and human authors.

The paper doesn't soft-pedal them, so neither will I:

Outside voices are interested but hedged — the right posture. Olivier Elemento (Weill Cornell): *"a real advance in terms of how we think about the publication process, with AI at the center."* Dongping Chen (U. Maryland) called executable papers *"quite compelling."* Nobody is claiming the peer-review replacement yet.

*Sources: Nature (Sep 16, 2026) — "Reimagining research papers as interactive and reliable AI agents" (Miao, Davis, Zhang, Pritchard, Zou); AI Weekly coverage (Sep 24, 2026); IEEE Spectrum; Stanford Medicine. Numbers as reported by the authors.*

**Companion notebook:** the runnable tutorial for this post — [download it here](https://danielsamfdo.github.io/blog/assets/paper2agent-mcp-demo.ipynb) (open in Colab/Jupyter).
