The first benchmark for professional work inside regulated biopharma.
We measure whether frontier agents can navigate contradictory records, identify controlling specifications, and deliver audit-ready professional work inside private biopharma companies.
Each environment places agents inside a company at a stage of bringing a medicine or medical test to patients. Its records and assignments reflect the work at that stage. Explore the environments below; Havenor and the Broad Institute include full walkthroughs.
1Discovery
Scientists find a promising molecule or idea.
2Lab & animal studies
Safety is tested in cells and animals before any person gets it.
3Human trials
Volunteers and patients receive it in carefully controlled studies.
GPT-6 Astra leads with a 70.4% mean score and fully passes 8 of 71 tasks. 6 of 9 models fully pass none.
Benchmark results for nine evaluated models
Rank
Model
Mean score ↓
Pass@1
Criteria passed
Cost
Steps
Tool calls
Time
1
GPT-6 AstraCodex · high
70.4%
8/71
542/752
$4.21
28.7
22.7
16.4 min
2
Claude Opus 5Claude Code · medium
63.8%
0/71
485/752
$2.68
29.2
29.4
10.6 min
3
Grok 4.6Cursor CLI · high
63.0%
2/71
480/752
$1.16
20.5
46.7
8.8 min
4
DeepSeek V4.1 FlashPi · high
57.0%
0/71
433/752
$0.11
36.2
40.5
6.3 min
5
GPT-6 SolCodex · high
54.3%
0/71
424/752
$0.64
28.2
26.3
5.8 min
6
Kimi K3Pi · max
48.5%
1/71
374/752
$0.72
21.9
25.1
4.9 min
7
GLM-5.3 FlashPi · high
47.6%
0/71
365/752
$0.04
22.8
25.8
3.1 min
8
Gemini 3.8 FlashCursor CLI · high
46.6%
0/71
356/752
$1.12
75.2
77.2
14.5 min
9
GPT-5.6 SolCodex · medium
46.6%
0/71
364/752
$1.27
29.2
23.2
4.1 min
Mean score: each task’s share of criteria passed, averaged over 71 tasks. Pass@1: tasks where every criterion passed in a single trial. Criteria passed is pooled across all 752 criteria. One scored trial per task, so results are subject to variance. Cost uses public API rates and excludes grading. Steps, tool calls, and time are per-task means.
Score against cost, effort, and release date
The dashed line joins models that no other model beats on both score and the chosen axis.
CostStepsTool callsTimeRelease date
GLM · Pi ★
DeepSeek · Pi ★
6 Sol · Codex
Kimi · Pi
Gemini · Cursor CLI
Grok · Cursor CLI ★
5.6 Sol · Codex
Opus · Claude Code ★
Astra · Codex ★
Each point represents one model. Cost, steps, tool calls, and agent execution time are per-task means across the same 71 tasks.
Two company environments are published in full, with documents, submitted files, and scoring checks. The trials shown here are examples; the leaderboard covers the full suite.
10 assignments at a fictional manufacturer.Partial-pilot examples are labeled.
4 assignments from a real FDA hold.Rebuilt verbatim from the Broad Institute's public IND record: the filed dossier, the agency's hold letter, and the lab data behind them. The best model passed most checks and still authorized work FDA had frozen.
A model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score.
All CapabilitiesHealth AuthoritiesDrug & Device ModalitiesDocument FormatsProfessional Families
Domain / SurfaceSpread & Range (Min → Max Macro Score)
Task: IND Partial Clinical Hold Response (MKL-P-01) · Environment: Broad Institute · FDA CDER IND 167326 Task prompt
"Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together."
What models did
Models frequently followed the internal steering memo’s proposal to "argue against the hold" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal.
Failed:Claude Opus 5, GPT-5.6 Sol, DeepSeek V4.1 Flash
What was required
Under 21 CFR 312.42(b)(1)(iv), the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order.
Passed:GPT-6 Astra Why it matters
If followed, healthy volunteers would have been enrolled while under an active federal clinical hold—a regulatory violation under 21 CFR 312.42 that subjects the presymptomatic trial to formal enforcement action. Submitted work
Open the files a model handed in and compare them with the starting files.
280 cells modified across 6 sheets (Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions).
Critical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045.
Showing sample modified cells before and after model execution:
Sheet: Control5 cell changes
Cell
Baseline State
Candidate Submission
C7
— (blank) 0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close.
C11
— (blank) Pull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet.
Sheet: Pull Plan18 cell changes
Cell
Baseline State
Candidate Submission
G2
— (blank) Hold - lot mapping correct, authorization withheld.
H2
— (blank) AVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete (QC-12) and the chamber qualified state is not established (QC-02).
Sheet: Regulatory Handoff36 cell changes
Cell
Baseline State
Candidate Submission
G2
— (blank) Not approved - scope and timing wrong.
H2
— (blank) HA-AVI401-IR-2026-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter.
What the benchmark covers
Regulators
US FDA
EU EMA
Health Canada
Korea MFDS
China NMPA / CDE
Products
mRNA vaccines
Antibody-drug conjugates
Targeted small molecules / kinase inhibitors
Peptide injectables
IVD and medical devices
Oral fixed-dose combinations
Parenteral biologics and MSAT
Documents
Common Technical Document (CTD) GxP quality records
Clinical and regulatory records
Statutory prescribing information
Jobs
Regulatory strategy and filing
Chemistry, manufacturing, and controls
Labeling and posology negotiation
Quality management and inspection readiness
Clinical trial oversight
How it works
What does a model get?
A simulated company with files, email, specifications, and business systems. The model takes an employee’s role and must hand in a specific document or spreadsheet.
Can models see future records?
No. Records stop at the task date. Havenor’s records end around September 2026, so tests scheduled after that have no results yet.
How is work scored?
Each submission is checked against a list of criteria for its task. Runs where grading failed for a technical reason are left out of the averages.
What is published?
Company documents, submitted files, scores, and summaries of what the model did. Full transcripts are available in Expert Review after signing in.
Can I run the tasks here?
These pages show completed trials. To run the full benchmark, including records, LIMS, and the verifier, use Harbor.
Why are the examples from Havenor Therapeutics?
Havenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page lists health authorities and product types, not individual companies.