cd /news/artificial-intelligence/expert-validated-stem-qa · home topics artificial-intelligence article
[ARTICLE · art-117347] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Expert-validated STEM QA

Researchers introduced 'Expert-validated STEM QA', a dataset of 398 expert-created questions in physics, chemistry, biology, and mathematics, built by 241 domain experts to address gaps in existing STEM benchmarks. Frontier AI models scored below 25% on the dataset, while post-training on a private 2,000-question version improved an open-source model's performance by 15% relative to baseline on the HLE-verified STEM subset (p=0.045). The team open-sourced part of the dataset for the AI research community.

read1 min views1 publishedSep 1, 2026

arXiv:2608.28591v1 Announce Type: new Abstract: Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present 'Expert-validated STEM QA', a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @expert-validated stem qa 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/expert-validated-ste…] indexed:0 read:1min 2026-09-01 ·