{"slug": "deepseek-beats-gpt-6-sol-in-autonomous-drug-development", "title": "DeepSeek beats GPT-6 Sol in autonomous drug development", "summary": "GPT-6 Astra led a new 71-task biopharma benchmark with a 70.4% mean score and fully passed 8 tasks, while GPT-6 Sol placed fifth at 54.3% and DeepSeek V4.1 Flash fourth at 57.0%, according to results published for nine evaluated models. Six of the nine models fully passed none of the 71 tasks, and on the Broad Institute's FDA CDER IND 167326 partial clinical hold task, Claude Opus 5, GPT-5.6 Sol and DeepSeek V4.1 Flash failed by following an internal memo's advice to argue against the hold and prepare trial sites in parallel, despite 21 CFR 312.42(b)(1)(iv) permitting only the symptomatic study.", "body_md": "The first benchmark for professional work inside regulated biopharma.\n\nWe measure whether frontier agents can navigate contradictory records, identify controlling specifications, and deliver audit-ready professional work inside private biopharma companies.\n\nEach environment places agents inside a company at a stage of bringing a medicine or medical test to patients. Its records and assignments reflect the work at that stage. Explore the environments below; Havenor and the Broad Institute include full walkthroughs.\n\n1Discovery\n\nScientists find a promising molecule or idea.\n\n2Lab & animal studies\n\nSafety is tested in cells and animals before any person gets it.\n\n3Human trials\n\nVolunteers and patients receive it in carefully controlled studies.\n\nGPT-6 Astra leads with a 70.4% mean score and fully passes 8 of 71 tasks. 6 of 9 models fully pass none.\n\nBenchmark results for nine evaluated models\n\nRank\n\nModel\n\nMean score ↓\n\nPass@1\n\nCriteria passed\n\nCost\n\nSteps\n\nTool calls\n\nTime\n\n1\n\nGPT-6 AstraCodex · high\n\n70.4%\n\n8/71\n\n542/752\n\n$4.21\n\n28.7\n\n22.7\n\n16.4 min\n\n2\n\nClaude Opus 5Claude Code · medium\n\n63.8%\n\n0/71\n\n485/752\n\n$2.68\n\n29.2\n\n29.4\n\n10.6 min\n\n3\n\nGrok 4.6Cursor CLI · high\n\n63.0%\n\n2/71\n\n480/752\n\n$1.16\n\n20.5\n\n46.7\n\n8.8 min\n\n4\n\nDeepSeek V4.1 FlashPi · high\n\n57.0%\n\n0/71\n\n433/752\n\n$0.11\n\n36.2\n\n40.5\n\n6.3 min\n\n5\n\nGPT-6 SolCodex · high\n\n54.3%\n\n0/71\n\n424/752\n\n$0.64\n\n28.2\n\n26.3\n\n5.8 min\n\n6\n\nKimi K3Pi · max\n\n48.5%\n\n1/71\n\n374/752\n\n$0.72\n\n21.9\n\n25.1\n\n4.9 min\n\n7\n\nGLM-5.3 FlashPi · high\n\n47.6%\n\n0/71\n\n365/752\n\n$0.04\n\n22.8\n\n25.8\n\n3.1 min\n\n8\n\nGemini 3.8 FlashCursor CLI · high\n\n46.6%\n\n0/71\n\n356/752\n\n$1.12\n\n75.2\n\n77.2\n\n14.5 min\n\n9\n\nGPT-5.6 SolCodex · medium\n\n46.6%\n\n0/71\n\n364/752\n\n$1.27\n\n29.2\n\n23.2\n\n4.1 min\n\nMean score: each task’s share of criteria passed, averaged over 71 tasks. Pass@1: tasks where every criterion passed in a single trial. Criteria passed is pooled across all 752 criteria. One scored trial per task, so results are subject to variance. Cost uses public API rates and excludes grading. Steps, tool calls, and time are per-task means.\n\nScore against cost, effort, and release date\n\nThe dashed line joins models that no other model beats on both score and the chosen axis.\n\nCostStepsTool callsTimeRelease date\n\nGLM · Pi ★\n\nDeepSeek · Pi ★\n\n6 Sol · Codex\n\nKimi · Pi\n\nGemini · Cursor CLI\n\nGrok · Cursor CLI ★\n\n5.6 Sol · Codex\n\nOpus · Claude Code ★\n\nAstra · Codex ★\n\nEach point represents one model. Cost, steps, tool calls, and agent execution time are per-task means across the same 71 tasks.\n\nTwo company environments are published in full, with documents, submitted files, and scoring checks. The trials shown here are examples; the leaderboard covers the full suite.\n\n10 assignments at a fictional manufacturer.Partial-pilot examples are labeled.\n\n4 assignments from a real FDA hold.Rebuilt verbatim from the Broad Institute's public IND record: the filed dossier, the agency's hold letter, and the lab data behind them. The best model passed most checks and still authorized work FDA had frozen.\n\nA model that leads on one kind of filing can trail on another. Filter the domains. Selecting a model on the leaderboard marks the rows where it is the high or low score.\n\nAll CapabilitiesHealth AuthoritiesDrug & Device ModalitiesDocument FormatsProfessional Families\n\nDomain / SurfaceSpread & Range (Min → Max Macro Score)\n\nTask: IND Partial Clinical Hold Response (MKL-P-01) · Environment: Broad Institute · FDA CDER IND 167326\n\nTask prompt\n\n\"Write the binding plan for the response: what may continue, what must stop, and how the answer to the FDA is put together.\"\n\nWhat models did\n\nModels frequently followed the internal steering memo’s proposal to \"argue against the hold\" and prepare trial sites in parallel, advising that screening and site setup could proceed quietly while drafting the rebuttal.\n\nFailed:Claude Opus 5, GPT-5.6 Sol, DeepSeek V4.1 Flash\n\nWhat was required\n\nUnder 21 CFR 312.42(b)(1)(iv), the symptomatic study was permitted, but the healthy-carrier study was on binding partial clinical hold. No screening, no site setup, and no volunteer contact was legally permissible until FDA formally issues a written removal order.\n\nPassed:GPT-6 Astra\n\nWhy it matters\n\nIf followed, healthy volunteers would have been enrolled while under an active federal clinical hold—a regulatory violation under 21 CFR 312.42 that subjects the presymptomatic trial to formal enforcement action.\n\nSubmitted work\n\nOpen the files a model handed in and compare them with the starting files.\n\n280 cells modified across 6 sheets (Control, Pull Plan, Execution Sequence, Chain and Controls, Regulatory Handoff, Review Actions).\n\nCritical Rubric Check: Correctly identifies that AVI401-R01 is 100 mL, but fails on qualified-state basis and misses chamber requalification lapse under AVT-SOP-045.\n\nShowing sample modified cells before and after model execution:\n\nSheet: Control5 cell changes\n\nCell\n\nBaseline State\n\nCandidate Submission\n\nC7\n\n— (blank)\n\n0.4 reviewed; not approved. Return as 0.5 after QC-01 to QC-14 close.\n\nC11\n\n— (blank)\n\nPull authorization, dataset release and regulatory handoff withheld. See Review Actions QC-01 to QC-15 and the QC Disposition sheet.\n\nSheet: Pull Plan18 cell changes\n\nCell\n\nBaseline State\n\nCandidate Submission\n\nG2\n\n— (blank)\n\nHold - lot mapping correct, authorization withheld.\n\nH2\n\n— (blank)\n\nAVI401-R01 is correctly listed as 100 mL. The pull is not authorized because the testing-chain clearance is not complete (QC-12) and the chamber qualified state is not established (QC-02).\n\nSheet: Regulatory Handoff36 cell changes\n\nCell\n\nBaseline State\n\nCandidate Submission\n\nG2\n\n— (blank)\n\nNot approved - scope and timing wrong.\n\nH2\n\n— (blank)\n\nHA-AVI401-IR-2026-014 point 7 states the 24-month pull is not part of sequence 0003 and is to be provided under COM-AVI401-008; no earlier pull is requested for the letter.\n\nWhat the benchmark covers\n\nRegulators\n\nUS FDA\n\nEU EMA\n\nHealth Canada\n\nKorea MFDS\n\nChina NMPA / CDE\n\nProducts\n\nmRNA vaccines\n\nAntibody-drug conjugates\n\nTargeted small molecules / kinase inhibitors\n\nPeptide injectables\n\nIVD and medical devices\n\nOral fixed-dose combinations\n\nParenteral biologics and MSAT\n\nDocuments\n\nCommon Technical Document (CTD)\n\nGxP quality records\n\nClinical and regulatory records\n\nStatutory prescribing information\n\nJobs\n\nRegulatory strategy and filing\n\nChemistry, manufacturing, and controls\n\nLabeling and posology negotiation\n\nQuality management and inspection readiness\n\nClinical trial oversight\n\nHow it works\n\nWhat does a model get?\n\nA simulated company with files, email, specifications, and business systems. The model takes an employee’s role and must hand in a specific document or spreadsheet.\n\nCan models see future records?\n\nNo. Records stop at the task date. Havenor’s records end around September 2026, so tests scheduled after that have no results yet.\n\nHow is work scored?\n\nEach submission is checked against a list of criteria for its task. Runs where grading failed for a technical reason are left out of the averages.\n\nWhat is published?\n\nCompany documents, submitted files, scores, and summaries of what the model did. Full transcripts are available in Expert Review after signing in.\n\nCan I run the tasks here?\n\nThese pages show completed trials. To run the full benchmark, including records, LIMS, and the verifier, use Harbor.\n\nWhy are the examples from Havenor Therapeutics?\n\nHavenor Therapeutics uses fictional company records that we can publish. For the rest of the suite, this page lists health authorities and product types, not individual companies.", "url": "https://wpnews.pro/news/deepseek-beats-gpt-6-sol-in-autonomous-drug-development", "canonical_source": "https://eval.raycaster.ai/benchmarks/biopharma-bench/", "published_at": "2026-09-25 15:23:03+00:00", "updated_at": "2026-09-25 15:33:00.958040+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "large-language-models", "ai-agents"], "entities": ["GPT-6 Astra", "GPT-6 Sol", "DeepSeek V4.1 Flash", "Claude Opus 5", "Broad Institute", "FDA", "Codex", "Pi"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/deepseek-beats-gpt-6-sol-in-autonomous-drug-development", "markdown": "https://wpnews.pro/news/deepseek-beats-gpt-6-sol-in-autonomous-drug-development.md", "text": "https://wpnews.pro/news/deepseek-beats-gpt-6-sol-in-autonomous-drug-development.txt", "jsonld": "https://wpnews.pro/news/deepseek-beats-gpt-6-sol-in-autonomous-drug-development.jsonld"}}