cd /news/artificial-intelligence/quality-diversity-stress-tests-for-p… · home topics artificial-intelligence article
[ARTICLE · art-91490] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

A new arXiv study (2608.08008v1) proposes quality-diversity stress testing for process reward models (PRMs) using MAP-Elites, revealing that Qwen2.5-Math-PRM-7B exhibits an aggregation-dependent vulnerability where padding yields 44 strict exploits with maximum gain 0.294 under mean pooling versus one exploit under minimum readout. The authors show that archive coverage alone cannot certify worst-case residual risk, but a predeclared paired LoRA repair protocol reduces exploit rates from 0.148 to 0.037 to 0.074 and lowers the worst attack from 0.333 to 0.177 to 0.212, improving ranking AUROC without degrading best-of-4 accuracy.

read1 min views1 publishedAug 11, 2026

arXiv:2608.08008v1 Announce Type: new Abstract: Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning. We formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctness-flipping edit in each behavior-space region while separating search coverage from exploit coverage. We characterize what such archives certify: finite-cell repair bounds covered-cell tail risk and average residual severity but cannot bound the worst remaining cell from covered fraction alone; under Lipschitz post-repair loss and metric-cover auditing, the residual is bounded by archive fitting error plus the Lipschitz constant times the covering radius. A controlled landscape validates this certificate and the impossibility of any fraction-only worst-case guarantee. On real PRMs, the search reveals an aggregation-dependent vulnerability in Qwen2.5-Math-PRM-7B: padding yields 44 strict exploits with maximum gain 0.294 under mean pooling versus one exploit under minimum readout; a matched syntactic control isolates the mechanism, and an RLHFlow value-head model shows the same qualitative effect with maximum gain 0.005. A predeclared paired LoRA repair protocol reduces exploit rates from 0.148 to 0.037 to 0.074, lowers the worst attack from 0.333 to 0.177 to 0.212, improves ranking AUROC without degrading best-of-4 accuracy, attributes gains to adversarial fine-tuning rather than archive diversity, and is confirmed by independent unpaired replications (44 to 1, clean-split worst gain 0.0092, MATH-500 41 to 0, clean ranking 40/40).

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/quality-diversity-st…] indexed:0 read:1min 2026-08-11 ·