TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On Researchers introduced TryOnReward, a fine-grained reward model for virtual try-on (VTON) built on a vision-language backbone that uses a foveation calibration objective and jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision. The team also built TryOnReward-100K, a human-annotated per-dimension rating dataset, plus the TryOn-Bench and TryOnRewardBench benchmarks, and reported that TryOnReward significantly outperforms generic judges in human preference alignment and yields human-preferred try-on results across multiple baselines when used as the reinforcement fine-tuning reward function. arXiv:2609.13259v1 Announce Type: new Abstract: Virtual Try-On VTON aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. However, classic fidelity metrics exhibit weak correlation with human judgments, and generic VLMs fail to provide the discriminative granularity demanded by try-on quality evaluation, which hinges on faithfully preserving garment and person details. This shortcoming is further exacerbated in the reinforcement fine-tuning RFT optimization and leads to severe reward hacking. To this end, we present TryOnReward, a fine-grained reward model tailored for VTON. Built on a vision-language backbone, it adopts a foveation calibration objective that grounds each quality dimension in the relevant region to avoid global shortcut learning. Meanwhile, TryOnReward jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision, leveraging both relative and absolute quality signals. For model training and evaluation, we build TryOnReward-100K, a human-annotated per-dimension rating dataset, alongside TryOn-Bench and TryOnRewardBench, two benchmarks covering diverse real scenarios. Extensive experiments confirm that TryOnReward significantly outperforms generic judges in human preference alignment, and when serving as the RFT reward function, it consistently yields human-preferred try-on results across multiple baselines.