cd /news/artificial-intelligence/labbench-can-ai-agents-decide-what-e… · home › topics › artificial-intelligence › article
[ARTICLE · art-142159] src=gamowlabs.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

LabBench: Can AI agents decide what experiment to run next?

Gamow Labs released LabBench, a 20-task benchmark built from real wet-lab records in drug discovery and genomics, finding that five frontier AI agents passed 182 of 406 grading criteria on no attempt and failed all 13 criteria on which experiment should come first. GPT-6 Astra and Claude Opus 5.5 tied on overall score, but Astra took a median 5 minutes per task versus Opus's 35, and criteria requiring choosing, committing or ranking passed 21% of the time against 47% for identifying what something is. On five core decisions no agent passed, appending a single sentence pointing at evidence GPT-6 Astra already had, without stating the answer, moved one task from 4/20 to 15/20, which Gamow Labs attributes to a recall rather than knowledge gap.

read3 min views1 publishedSep 30, 2026
LabBench: Can AI agents decide what experiment to run next?
Image: source

Accelerating biological experimentation relies on quickly deciding which experiment comes next. Picking the test that can fastest invalidate or support a hypothesis is a skill only human experts perform reliably today. To scale discovery past human capability, we have to evaluate and improve AI agents on it.

LabBench is 20 held-out tasks built from real wet-lab records in drug discovery and genomics. Each is a snapshot in time: the agent gets the records that existed at a decision point, and the lab’s interpretation and decision are withheld. It must commit to the next step, graded against 20–22 binary criteria tied to what the lab actually decided.

The tasks come from real lab data unlikely to be recalled from the open internet. We expect the agent to navigate the evidence as a real researcher would, without the brief pointing to what matters.

Results #

Five frontier agents each ran once per task in their vendor’s harness.

GPT-6 Astra and Claude Opus 5.5 tie, but Astra took a median 5 minutes per task to Opus’s 35.

Agents fail the same criteria

182 of 406 criteria were passed by no agent; 31 by all five. What one frontier agent misses, the others usually miss too.

They interpret; they don’t choose

Agents excel at interpreting previous experiments: saying what a measurement is, declining an overclaim, reconstructing an analysis. Criteria that require choosing, committing or ranking pass 21% of the time, against 47% for identifying what something is. No agent passed any of the 13 criteria on which experiment should come first.

Where agents differ

Astra leads on core decisions (52%), Opus on evidence integration (53%). On experiment design, the best agent passes 9%.

The knowledge is latent #

A common issue with hard benchmarks is unreasonable criteria that make tasks effectively impossible. We tested this directly. On five core decisions no agent passed, we appended one sentence pointing at evidence GPT-6 Astra already had, without stating the answer. It passed all five.

Drug response

Evidence

A compound triggers the strongest early NF-κB signature of four death inducers; nothing tests whether it drives death.

Agents

Four of five designed the NF-κB test, then ranked it second.

Hint

“Rank first the experiment that could tell a cause from a bystander.” Score: 4/20 → 15/20.

Assay development

Agents

All five chose the larger 238 bp peak and ignored the shoulder of uncut DNA.

Hint

“Judge each tagmentation trace by its whole size distribution…” The agent chose the lab’s condition.

The hints add no biology; they redirect attention. The models have the knowledge but do not reliably recall it.

A reasoning bottleneck #

Frontier agents know enough biology to work alongside expert biologists when experimental design stays with humans. Their ability to design experiments independently is weak: they fail to commit to an experiment and to discriminate the best one from the alternatives.

In real biological experimentation, wall-clock time is irreducible. Cells grow, differentiate and respond on their own schedules, and each experiment consumes weeks and material. Autonomous and correct experimental design is arguably the single most important capability for making AI agents superhuman at these tasks.

LabBench measures this directly. Improved results should indicate agents that can carry out supervised, autonomous wet-lab experimentation, and eventually run full programs themselves. To get there, models must develop a more cohesive world model of what experiments cost and of the specific evidence that would count against a hypothesis.

Work with us #

Gamow Labs is uniquely equipped to produce data at the frontier of AI and biological experimentation. We are capable of producing the aforementioned tasks at scale. Reach out.

Full methods are in the preprint. We’re also hiring.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @gamow labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/labbench-can-ai-agen…] indexed:0 read:3min 2026-09-30 · —