{"slug": "labbench-can-ai-agents-decide-what-experiment-to-run-next", "title": "LabBench: Can AI agents decide what experiment to run next?", "summary": "Gamow Labs released LabBench, a 20-task benchmark built from real wet-lab records in drug discovery and genomics, finding that five frontier AI agents passed 182 of 406 grading criteria on no attempt and failed all 13 criteria on which experiment should come first. GPT-6 Astra and Claude Opus 5.5 tied on overall score, but Astra took a median 5 minutes per task versus Opus's 35, and criteria requiring choosing, committing or ranking passed 21% of the time against 47% for identifying what something is. On five core decisions no agent passed, appending a single sentence pointing at evidence GPT-6 Astra already had, without stating the answer, moved one task from 4/20 to 15/20, which Gamow Labs attributes to a recall rather than knowledge gap.", "body_md": "Accelerating biological experimentation relies on quickly deciding which experiment comes next. Picking the test that can fastest invalidate or support a hypothesis is a skill only human experts perform reliably today. To scale discovery past human capability, we have to evaluate and improve AI agents on it.\n\n**LabBench** is 20 held-out tasks built from real wet-lab records in drug discovery and genomics. Each is a snapshot in time: the agent gets the records that existed at a decision point, and the lab’s interpretation and decision are withheld. It must commit to the next step, graded against 20–22 binary criteria tied to what the lab actually decided.\n\nThe tasks come from real lab data unlikely to be recalled from the open internet. We expect the agent to navigate the evidence as a real researcher would, without the brief pointing to what matters.\n\n## Results\n\nFive frontier agents each ran once per task in their vendor’s harness.\n\nGPT-6 Astra and Claude Opus 5.5 tie, but Astra took a median 5 minutes per task to Opus’s 35.\n\n### Agents fail the same criteria\n\n182 of 406 criteria were passed by no agent; 31 by all five. What one frontier agent misses, the others usually miss too.\n\n### They interpret; they don’t choose\n\nAgents excel at interpreting previous experiments: saying what a measurement is, declining an overclaim, reconstructing an analysis. Criteria that require choosing, committing or ranking pass 21% of the time, against 47% for identifying what something is. No agent passed any of the 13 criteria on which experiment should come first.\n\n### Where agents differ\n\nAstra leads on core decisions (52%), Opus on evidence integration (53%). On experiment design, the best agent passes 9%.\n\n## The knowledge is latent\n\nA common issue with hard benchmarks is unreasonable criteria that make tasks effectively impossible. We tested this directly. On five core decisions no agent passed, we appended one sentence pointing at evidence GPT-6 Astra already had, without stating the answer. It passed all five.\n\n#### Drug response\n\nEvidence\n\nA compound triggers the strongest early NF-κB signature of four death inducers; nothing tests whether it drives death.\n\nAgents\n\nFour of five designed the NF-κB test, then ranked it second.\n\nHint\n\n“Rank first the experiment that could tell a cause from a bystander.” Score: 4/20 → 15/20.\n\n#### Assay development\n\nAgents\n\nAll five chose the larger 238 bp peak and ignored the shoulder of uncut DNA.\n\nHint\n\n“Judge each tagmentation trace by its whole size distribution…” The agent chose the lab’s condition.\n\nThe hints add no biology; they redirect attention. The models have the knowledge but do not reliably recall it.\n\n## A reasoning bottleneck\n\nFrontier agents know enough biology to work alongside expert biologists when experimental design stays with humans. Their ability to design experiments independently is weak: they fail to commit to an experiment and to discriminate the best one from the alternatives.\n\nIn real biological experimentation, wall-clock time is irreducible. Cells grow, differentiate and respond on their own schedules, and each experiment consumes weeks and material. Autonomous and correct experimental design is arguably the single most important capability for making AI agents superhuman at these tasks.\n\nLabBench measures this directly. Improved results should indicate agents that can carry out supervised, autonomous wet-lab experimentation, and eventually run full programs themselves. To get there, models must develop a more cohesive world model of what experiments cost and of the specific evidence that would count against a hypothesis.\n\n## Work with us\n\nGamow Labs is uniquely equipped to produce data at the frontier of AI and biological experimentation. We are capable of producing the aforementioned tasks at scale. Reach out.\n\nFull methods are in the [preprint](assets/labbench-preprint.pdf). We’re also hiring.", "url": "https://wpnews.pro/news/labbench-can-ai-agents-decide-what-experiment-to-run-next", "canonical_source": "https://gamowlabs.com/labbench-benchmarking-ai-wet-lab-decisions.html", "published_at": "2026-09-30 00:59:52+00:00", "updated_at": "2026-09-30 01:17:32.593642+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "ai-safety"], "entities": ["Gamow Labs", "LabBench", "GPT-6 Astra", "Claude Opus 5.5", "NF-κB"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/labbench-can-ai-agents-decide-what-experiment-to-run-next", "markdown": "https://wpnews.pro/news/labbench-can-ai-agents-decide-what-experiment-to-run-next.md", "text": "https://wpnews.pro/news/labbench-can-ai-agents-decide-what-experiment-to-run-next.txt", "jsonld": "https://wpnews.pro/news/labbench-can-ai-agents-decide-what-experiment-to-run-next.jsonld"}}