{"slug": "renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability", "title": "ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration", "summary": "Researchers propose ReNFT, a method that repairs mode collapse in reward post-training of diffusion generators by recalibrating internal probability mass, preserving 98.9% and 99.0% of NFT's reward on PickScore and GenEval while improving DreamSim-Div by 58.8% and 55.0%, respectively. The approach uses unconditional probes and paired counterfactual proposals to reverse collapse without external signals, offering a complementary alternative to existing interventions.", "body_md": "arXiv:2609.00061v1 Announce Type: new\nAbstract: Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learning new visual content. Collapse is therefore suppression, not deletion, and can be reversed from within the generator. We propose ReNFT, which repairs a high-reward, low-diversity adapter through internal probability-mass recalibration. Unconditional probes first prioritize \"anti-hub\" prompts where the prompt-independent bias is easiest to expose. Two policy-dominated mixed routes then generate matched counterfactual proposals from the same prompt and initial noise, one probing the frozen base direction for suppressed alternatives and the other exposing the post-trained unconditional tendency. Reward ranking with an adaptive flipping guard assigns pull and push roles, and a joint-and-paired NFT update realizes the repair. On PickScore and GenEval, ReNFT retains 98.9% and 99.0% of NFT's reward while improving DreamSim-Div by 58.8% and 55.0%, respectively, offering a complementary alternative to external interventions.", "url": "https://wpnews.pro/news/renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability", "canonical_source": "https://arxiv.org/abs/2609.00061", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:23:48.390446+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research"], "entities": ["ReNFT", "PickScore", "GenEval", "DreamSim-Div", "NFT"], "alternates": {"html": "https://wpnews.pro/news/renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability", "markdown": "https://wpnews.pro/news/renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability.md", "text": "https://wpnews.pro/news/renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability.txt", "jsonld": "https://wpnews.pro/news/renft-repairing-mode-collapse-in-reward-post-training-via-internal-probability.jsonld"}}