cd /news/ai-research/opendiscoverytrace-process-traces-fo… · home topics ai-research article
[ARTICLE · art-126520] src=arxiv.org ↗ pub= topic=ai-research verified=true sentiment=· neutral

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

Researchers released OpenDiscoveryTrace, a public dataset of 558 complete AI scientific agent trajectories across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis, spanning seven models including GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Pilot analysis of 363 LLM-judged trajectories found all three frontier models scored comparable success rates of 84–89%, but Claude Opus 4.6 produced 30 times more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, p < 0.0001, Cliff's δ = 0.613), with 66.7% of Claude's errors being tool misuse versus 83.6% reasoning errors for GPT-5.4. The dataset, trace schema, agent harness, and five benchmark tasks with logistic regression, random forest, LSTM, and Transformer baselines are released under CC BY 4.0 to support process-level evaluation, scientific agent auditing, and AI governance.

by read1 min views3 publishedSep 11, 2026

arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis. The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories. Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation: all three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30$\times$ more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, $p < 0.0001$, Cliff's $\delta = 0.613$), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4. We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent harness, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.

── more in #ai-research 4 stories · sorted by recency
── more on @opendiscoverytrace 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/opendiscoverytrace-p…] indexed:0 read:1min 2026-09-11 ·