cd /news/large-language-models/more-correct-mass-worse-answers-why-… · home topics large-language-models article
[ARTICLE · art-99419] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

Researchers at arXiv (paper 2608.14420v1) found that Power Sampling, a method that sharpens a language model's distribution over complete generation trajectories, can increase probability mass on correct answers while degrading downstream inference, causing accuracy drops of up to 18.5 percentage points across models and reasoning benchmarks. The team traced the issue to dose and coverage mismatches and proposed a deformation-controlled, support-preserving Power target that reverses the losses and outperforms standard multi-sample inference in same-budget tests with weighted self-consistency.

read1 min views2 publishedAug 17, 2026

arXiv:2608.14420v1 Announce Type: new Abstract: Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. However, we uncover a striking paradox: Power Sampling can drive more probability mass toward correct trajectories while degrading the downstream inference it is intended to enhance. Using self-consistency as a representative case, we observe accuracy drops of up to 18.5 percentage points across models and reasoning benchmarks. We trace this paradox to two mismatches. Dose mismatch arises because a fixed exponent induces drastically different amounts of distributional change across problems. Coverage mismatch arises because global sharpening concentrates mass on a narrow set of dominant paths: high pass@k, often interpreted as evidence of preserved diversity, can therefore coexist with the loss of broad reasoning-path support required for downstream aggregation, search, and selection. Guided by this diagnosis, we replace uniform trajectory exponentiation with a deformation-controlled, support-preserving Power target that calibrates sharpening across problems while limiting the suppression of moderate-probability paths. In a same-budget instantiation with weighted self-consistency, the repaired sampler reverses the losses caused by global Power and outperforms standard multi-sample inference across reasoning benchmarks.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/more-correct-mass-wo…] indexed:0 read:1min 2026-08-17 ·