{"slug": "chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation", "title": "CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation", "summary": "Researchers introduced CHORUS, a post-training framework that combines behaviorally diverse LLM checkpoints into a single 4B model, achieving 88.0% Pass@1 on the CVDP-ECov hardware verification benchmark, outperforming DeepSeek-R1 (671B) by 13.5 percentage points. The framework leverages staged supervised fine-tuning and dense-reward reinforcement learning to create complementary experts, which are then merged or further post-trained to exceed the best individual expert's performance.", "body_md": "arXiv:2608.10090v1 Announce Type: new\nAbstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.", "url": "https://wpnews.pro/news/chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation", "canonical_source": "https://arxiv.org/abs/2608.10090", "published_at": "2026-08-12 04:00:00+00:00", "updated_at": "2026-08-12 04:17:55.675838+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["CHORUS", "DeepSeek-R1", "CVDP-ECov"], "alternates": {"html": "https://wpnews.pro/news/chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation", "markdown": "https://wpnews.pro/news/chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation.md", "text": "https://wpnews.pro/news/chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation.txt", "jsonld": "https://wpnews.pro/news/chorus-complementary-experts-for-high-coverage-testbench-stimulus-generation.jsonld"}}