cd /news/computer-vision/tryonreward-learning-foveated-consis… · home topics computer-vision article
[ARTICLE · art-129858] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On

Researchers introduced TryOnReward, a fine-grained reward model for virtual try-on (VTON) built on a vision-language backbone that uses a foveation calibration objective and jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision. The team also built TryOnReward-100K, a human-annotated per-dimension rating dataset, plus the TryOn-Bench and TryOnRewardBench benchmarks, and reported that TryOnReward significantly outperforms generic judges in human preference alignment and yields human-preferred try-on results across multiple baselines when used as the reinforcement fine-tuning reward function.

by read1 min views1 publishedSep 15, 2026

arXiv:2609.13259v1 Announce Type: new Abstract: Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. However, classic fidelity metrics exhibit weak correlation with human judgments, and generic VLMs fail to provide the discriminative granularity demanded by try-on quality evaluation, which hinges on faithfully preserving garment and person details. This shortcoming is further exacerbated in the reinforcement fine-tuning (RFT) optimization and leads to severe reward hacking. To this end, we present TryOnReward, a fine-grained reward model tailored for VTON. Built on a vision-language backbone, it adopts a foveation calibration objective that grounds each quality dimension in the relevant region to avoid global shortcut learning. Meanwhile, TryOnReward jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision, leveraging both relative and absolute quality signals. For model training and evaluation, we build TryOnReward-100K, a human-annotated per-dimension rating dataset, alongside TryOn-Bench and TryOnRewardBench, two benchmarks covering diverse real scenarios. Extensive experiments confirm that TryOnReward significantly outperforms generic judges in human preference alignment, and when serving as the RFT reward function, it consistently yields human-preferred try-on results across multiple baselines.

── more in #computer-vision 4 stories · sorted by recency
── more on @tryonreward 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tryonreward-learning…] indexed:0 read:1min 2026-09-15 ·