AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

arXiv · AI, language, vision and robotics · article · Aug 30, 2026 · UTC

Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learni

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:31:56.984Z. This is not the publication date.