cd /news/large-language-models/qvac-genesis-iii-a-large-scale-high-… · home topics large-language-models article
[ARTICLE · art-133316] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

Researchers introduced QVAC Genesis III, a 191.43B-token STEM-focused synthetic corpus spanning 19 domains, built via a dual generation strategy that uses a weak edge-scale student model's failures as corrective explanations and its successes as contrastive option-level reasoning. In controlled from-scratch ablations with 1.7B-parameter models, QVAC Genesis III-trained models outperformed both Cosmopedia-v2-trained models and the publicly released Cosmo-1B across ARC, GPQA Diamond, and MMLU STEM benchmarks, gaining up to +28.57% on ARC-E and +21.35% on ARC-C and reaching a Valid Answer Rate of up to 99.45%. The corpus targets edge AI and on-device deployment where token budgets are tightly constrained.

by read1 min views1 publishedSep 18, 2026

arXiv:2609.19513v1 Announce Type: new Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy that performs targeted teacher distillation using a weak edge-scale student model as signal: the student's failures are converted into corrective explanations, while its successes are expanded into contrastive option-level reasoning over all answer choices. We further introduce an LLM-as-a-parser evaluation protocol that extracts final answers from free-form outputs and tracks both accuracy and answer validity. To validate the effectiveness of our QVAC Genesis III data, we conduct controlled from-scratch ablations with 1.7B-parameter models, showing that models trained with QVAC Genesis III consistently outperform both models trained with the open-source synthetic corpus Cosmopedia-v2 and the publicly released Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% on ARC-E and +21.35% on ARC-C, while reaching a Valid Answer Rate of up to 99.45%.

── more in #large-language-models 4 stories · sorted by recency
── more on @qvac genesis iii 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qvac-genesis-iii-a-l…] indexed:0 read:1min 2026-09-18 ·