cd /news/ai-safety/deepseek-beats-gpt-6-sol-in-autonomo… · home › topics › ai-safety › article
[ARTICLE · art-139725] src=eval.raycaster.ai ↗ pub= topic=ai-safety verified=true sentiment=· neutral

DeepSeek beats GPT-6 Sol in autonomous drug development

GPT-6 Astra led a new 71-task biopharma benchmark with a 70.4% mean score and fully passed 8 tasks, while GPT-6 Sol placed fifth at 54.3% and DeepSeek V4.1 Flash fourth at 57.0%, according to results published for nine evaluated models. Six of the nine models fully passed none of the 71 tasks, and on the Broad Institute's FDA CDER IND 167326 partial clinical hold task, Claude Opus 5, GPT-5.6 Sol and DeepSeek V4.1 Flash failed by following an internal memo's advice to argue against the hold and prepare trial sites in parallel, despite 21 CFR 312.42(b)(1)(iv) permitting only the symptomatic study.

read6 min views1 publishedSep 25, 2026
DeepSeek beats GPT-6 Sol in autonomous drug development
Image: source

The first benchmark for professional work inside regulated biopharma.

We measure whether frontier agents can navigate contradictory records, identify controlling specifications, and deliver audit-ready professional work inside private biopharma companies.

Each environment places agents inside a company at a stage of bringing a medicine or medical test to patients. Its records and assignments reflect the work at that stage. Explore the environments below; Havenor and the Broad Institute include full walkthroughs.

1Discovery

Scientists find a promising molecule or idea.

2Lab & animal studies

Safety is tested in cells and animals before any person gets it.

3Human trials

Volunteers and patients receive it in carefully controlled studies.

GPT-6 Astra leads with a 70.4% mean score and fully passes 8 of 71 tasks. 6 of 9 models fully pass none.

Benchmark results for nine evaluated models

Rank

Model

Mean score ↓

Pass@1

Criteria passed

Cost

Steps

Tool calls

Time

1

GPT-6 AstraCodex · high

70.4%

8/71

542/752

$4.21

28.7

22.7

16.4 min

2

Claude Opus 5Claude Code · medium

63.8%

0/71

485/752

$2.68

29.2

29.4

10.6 min

3

Grok 4.6Cursor CLI · high

63.0%

2/71

480/752

$1.16

20.5

46.7

8.8 min

4

DeepSeek V4.1 FlashPi · high

57.0%

0/71

433/752

$0.11

36.2

40.5

6.3 min

5

GPT-6 SolCodex · high

54.3%

0/71

424/752

$0.64

28.2

26.3

5.8 min

6

Kimi K3Pi · max

48.5%

1/71

374/752

$0.72

21.9

25.1

4.9 min

7

GLM-5.3 FlashPi · high

47.6%

0/71

365/752

$0.04

22.8

25.8

3.1 min

8

Gemini 3.8 FlashCursor CLI · high

46.6%

0/71

356/752

$1.12

75.2

77.2

14.5 min

9

GPT-5.6 SolCodex · medium

46.6%

0/71

364/752

$1.27

29.2

23.2

4.1 min

Mean score: each task’s share of criteria passed, averaged over 71 tasks. Pass@1: tasks where every criterion passed in a single trial. Criteria passed is pooled across all 752 criteria. One scored trial per task, so results are subject to variance. Cost uses public API rates and excludes grading. Steps, tool calls, and time are per-task means.

Score against cost, effort, and release date

The dashed line joins models that no other model beats on both score and the chosen axis.

CostStepsTool callsTimeRelease date

GLM · Pi ★

DeepSeek · Pi ★

6 Sol · Codex

Kimi · Pi

Gemini · Cursor CLI

Grok · Cursor CLI ★

5.6 Sol · Codex

Opus · Claude Code ★

Astra · Codex ★

Each point represents one model. Cost, steps, tool calls, and agent execution time are per-task means across the same 71 tasks.

Two company environments are published in full, with documents, submitted files, and scoring checks. The trials shown here are examples; the leaderboard covers the full suite.

10 assignments at a fictional manufacturer.Partial-pilot examples are labeled.

4 assignments from a real FDA hold.Rebuilt verbatim from the Broad Institute's public IND record: the filed dossier, the agency's hold letter, and the lab data behind them. The best model passed most checks and still authorized work FDA had frozen.

A model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score.

All CapabilitiesHealth AuthoritiesDrug & Device ModalitiesDocument FormatsProfessional Families

Domain / SurfaceSpread & Range (Min → Max Macro Score)

Task: IND Partial Clinical Hold Response (MKL-P-01) · Environment: Broad Institute · FDA CDER IND 167326 Task prompt

"Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together."

What models did

Models frequently followed the internal steering memo’s proposal to "argue against the hold" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal.

Failed:Claude Opus 5, GPT-5.6 Sol, DeepSeek V4.1 Flash

What was required

Under 21 CFR 312.42(b)(1)(iv), the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order.

Passed:GPT-6 Astra Why it matters

If followed, healthy volunteers would have been enrolled while under an active federal clinical hold—a regulatory violation under 21 CFR 312.42 that subjects the presymptomatic trial to formal enforcement action. Submitted work

Open the files a model handed in and compare them with the starting files.

280 cells modified across 6 sheets (Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions).

Critical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045.

Showing sample modified cells before and after model execution:

Sheet: Control5 cell changes

Cell

Baseline State

Candidate Submission

C7

— (blank) 0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close.

C11

— (blank) Pull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet.

Sheet: Pull Plan18 cell changes

Cell

Baseline State

Candidate Submission

G2

— (blank) Hold - lot mapping correct, authorization withheld.

H2

— (blank) AVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete (QC-12) and the chamber qualified state is not established (QC-02).

Sheet: Regulatory Handoff36 cell changes

Cell

Baseline State

Candidate Submission

G2

— (blank) Not approved - scope and timing wrong.

H2

— (blank) HA-AVI401-IR-2026-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter.

What the benchmark covers

Regulators

US FDA

EU EMA

Health Canada

Korea MFDS

China NMPA / CDE

Products

mRNA vaccines

Antibody-drug conjugates

Targeted small molecules / kinase inhibitors

Peptide injectables

IVD and medical devices

Oral fixed-dose combinations

Parenteral biologics and MSAT

Documents

Common Technical Document (CTD) GxP quality records

Clinical and regulatory records

Statutory prescribing information

Jobs

Regulatory strategy and filing

Chemistry, manufacturing, and controls

Labeling and posology negotiation

Quality management and inspection readiness

Clinical trial oversight

How it works

What does a model get?

A simulated company with files, email, specifications, and business systems. The model takes an employee’s role and must hand in a specific document or spreadsheet.

Can models see future records?

No. Records stop at the task date. Havenor’s records end around September 2026, so tests scheduled after that have no results yet.

How is work scored?

Each submission is checked against a list of criteria for its task. Runs where grading failed for a technical reason are left out of the averages.

What is published?

Company documents, submitted files, scores, and summaries of what the model did. Full transcripts are available in Expert Review after signing in.

Can I run the tasks here?

These pages show completed trials. To run the full benchmark, including records, LIMS, and the verifier, use Harbor.

Why are the examples from Havenor Therapeutics?

Havenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page lists health authorities and product types, not individual companies.

── more in #ai-safety 4 stories · sorted by recency
── more on @gpt-6 astra 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-beats-gpt-6…] indexed:0 read:6min 2026-09-25 · —