DeepSeek beats GPT-6 Sol in autonomous drug development GPT-6 Astra led a new 71-task biopharma benchmark with a 70.4% mean score and fully passed 8 tasks, while GPT-6 Sol placed fifth at 54.3% and DeepSeek V4.1 Flash fourth at 57.0%, according to results published for nine evaluated models. Six of the nine models fully passed none of the 71 tasks, and on the Broad Institute's FDA CDER IND 167326 partial clinical hold task, Claude Opus 5, GPT-5.6 Sol and DeepSeek V4.1 Flash failed by following an internal memo's advice to argue against the hold and prepare trial sites in parallel, despite 21 CFR 312.42(b)(1)(iv) permitting only the symptomatic study. The first benchmark for professional work inside regulated biopharma. We measure whether frontier agents can navigate contradictory records, identify controlling specifications, and deliver audit-ready professional work inside private biopharma companies. Each environment places agents inside a company at a stage of bringing a medicine or medical test to patients. Its records and assignments reflect the work at that stage. Explore the environments below; Havenor and the Broad Institute include full walkthroughs. 1Discovery Scientists find a promising molecule or idea. 2Lab & animal studies Safety is tested in cells and animals before any person gets it. 3Human trials Volunteers and patients receive it in carefully controlled studies. GPT-6 Astra leads with a 70.4% mean score and fully passes 8 of 71 tasks. 6 of 9 models fully pass none. Benchmark results for nine evaluated models Rank Model Mean score ↓ Pass@1 Criteria passed Cost Steps Tool calls Time 1 GPT-6 AstraCodex · high 70.4% 8/71 542/752 $4.21 28.7 22.7 16.4 min 2 Claude Opus 5Claude Code · medium 63.8% 0/71 485/752 $2.68 29.2 29.4 10.6 min 3 Grok 4.6Cursor CLI · high 63.0% 2/71 480/752 $1.16 20.5 46.7 8.8 min 4 DeepSeek V4.1 FlashPi · high 57.0% 0/71 433/752 $0.11 36.2 40.5 6.3 min 5 GPT-6 SolCodex · high 54.3% 0/71 424/752 $0.64 28.2 26.3 5.8 min 6 Kimi K3Pi · max 48.5% 1/71 374/752 $0.72 21.9 25.1 4.9 min 7 GLM-5.3 FlashPi · high 47.6% 0/71 365/752 $0.04 22.8 25.8 3.1 min 8 Gemini 3.8 FlashCursor CLI · high 46.6% 0/71 356/752 $1.12 75.2 77.2 14.5 min 9 GPT-5.6 SolCodex · medium 46.6% 0/71 364/752 $1.27 29.2 23.2 4.1 min Mean score: each task’s share of criteria passed, averaged over 71 tasks. Pass@1: tasks where every criterion passed in a single trial. Criteria passed is pooled across all 752 criteria. One scored trial per task, so results are subject to variance. Cost uses public API rates and excludes grading. Steps, tool calls, and time are per-task means. Score against cost, effort, and release date The dashed line joins models that no other model beats on both score and the chosen axis. CostStepsTool callsTimeRelease date GLM · Pi ★ DeepSeek · Pi ★ 6 Sol · Codex Kimi · Pi Gemini · Cursor CLI Grok · Cursor CLI ★ 5.6 Sol · Codex Opus · Claude Code ★ Astra · Codex ★ Each point represents one model. Cost, steps, tool calls, and agent execution time are per-task means across the same 71 tasks. Two company environments are published in full, with documents, submitted files, and scoring checks. The trials shown here are examples; the leaderboard covers the full suite. 10 assignments at a fictional manufacturer.Partial-pilot examples are labeled. 4 assignments from a real FDA hold.Rebuilt verbatim from the Broad Institute's public IND record: the filed dossier, the agency's hold letter, and the lab data behind them. The best model passed most checks and still authorized work FDA had frozen. A model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score. All CapabilitiesHealth AuthoritiesDrug & Device ModalitiesDocument FormatsProfessional Families Domain / SurfaceSpread & Range Min → Max Macro Score Task: IND Partial Clinical Hold Response MKL-P-01 · Environment: Broad Institute · FDA CDER IND 167326 Task prompt "Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together." What models did Models frequently followed the internal steering memo’s proposal to "argue against the hold" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal. Failed:Claude Opus 5, GPT-5.6 Sol, DeepSeek V4.1 Flash What was required Under 21 CFR 312.42 b 1 iv , the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order. Passed:GPT-6 Astra Why it matters If followed, healthy volunteers would have been enrolled while under an active federal clinical hold—a regulatory violation under 21 CFR 312.42 that subjects the presymptomatic trial to formal enforcement action. Submitted work Open the files a model handed in and compare them with the starting files. 280 cells modified across 6 sheets Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions . Critical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045. Showing sample modified cells before and after model execution: Sheet: Control5 cell changes Cell Baseline State Candidate Submission C7 — blank 0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close. C11 — blank Pull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet. Sheet: Pull Plan18 cell changes Cell Baseline State Candidate Submission G2 — blank Hold - lot mapping correct, authorization withheld. H2 — blank AVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete QC-12 and the chamber qualified state is not established QC-02 . Sheet: Regulatory Handoff36 cell changes Cell Baseline State Candidate Submission G2 — blank Not approved - scope and timing wrong. H2 — blank HA-AVI401-IR-2026-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter. What the benchmark covers Regulators US FDA EU EMA Health Canada Korea MFDS China NMPA / CDE Products mRNA vaccines Antibody-drug conjugates Targeted small molecules / kinase inhibitors Peptide injectables IVD and medical devices Oral fixed-dose combinations Parenteral biologics and MSAT Documents Common Technical Document CTD GxP quality records Clinical and regulatory records Statutory prescribing information Jobs Regulatory strategy and filing Chemistry, manufacturing, and controls Labeling and posology negotiation Quality management and inspection readiness Clinical trial oversight How it works What does a model get? A simulated company with files, email, specifications, and business systems. The model takes an employee’s role and must hand in a specific document or spreadsheet. Can models see future records? No. Records stop at the task date. Havenor’s records end around September 2026, so tests scheduled after that have no results yet. How is work scored? Each submission is checked against a list of criteria for its task. Runs where grading failed for a technical reason are left out of the averages. What is published? Company documents, submitted files, scores, and summaries of what the model did. Full transcripts are available in Expert Review after signing in. Can I run the tasks here? These pages show completed trials. To run the full benchmark, including records, LIMS, and the verifier, use Harbor. Why are the examples from Havenor Therapeutics? Havenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page lists health authorities and product types, not individual companies.