cd /news/artificial-intelligence/ocgquant-outlier-companion-grouping-… · home topics artificial-intelligence article
[ARTICLE · art-118566] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

Researchers introduced OCGQuant, a post-training quantization method that pairs outlier channels with low-magnitude companions to reduce quantization error in NVFP4 low-bit inference, achieving the lowest WikiText-2 perplexity and highest average downstream accuracy on Llama3 and Qwen3 among evaluated methods while maintaining prefill speedup close to RTN and matching its peak decoding memory. The method, detailed in arXiv:2609.00066v1, is available at https://github.com/Eshamont/OCGQuant.

read1 min views2 publishedSep 2, 2026

arXiv:2609.00066v1 Announce Type: new Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the remaining values sharing the same scale. Existing post-training quantization (PTQ) methods mitigate outlier errors through strategies such as mixed precision, rotation, or residual compensation, but these approaches are either not specifically tailored to NVFP4 or introduce additional computation. In this work, we revisit NVFP4 from a channel-grouping perspective and define the reducible error incurred by remaining block values under the scale set by the block maximum as Collateral Quantization Error. Based on this insight, we propose OCGQuant, a post-training quantization method centered on Outlier-Companion Grouping (OCG), which adaptively pairs outlier channels with low-magnitude companion channels to improve NVFP4 activation block composition. Experiments on Llama3 and Qwen3 show that OCGQuant achieves the lowest WikiText-2 perplexity and highest average downstream accuracy among evaluated PTQ methods, while maintaining prefill speedup close to RTN and matching its peak decoding memory. Code is available at https://github.com/Eshamont/OCGQuant.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ocgquant 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ocgquant-outlier-com…] indexed:0 read:1min 2026-09-02 ·