{"slug": "comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable", "title": "COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping", "summary": "Researchers introduced COMEX, a benchmark with 33,161 quadruples of expanded images, crop boxes, composition categories, and explanations, to reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. They proposed a two-stage SFT+GRPO framework that improves crop quality, composition prediction, and explanation faithfulness, benchmarking 15 large vision-language models and existing cropping methods on COMEX.", "body_md": "arXiv:2608.07570v1 Announce Type: new\nAbstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.", "url": "https://wpnews.pro/news/comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable", "canonical_source": "https://arxiv.org/abs/2608.07570", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 04:24:25.617221+00:00", "lang": "en", "topics": ["computer-vision", "large-language-models", "generative-ai"], "entities": ["COMEX", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable", "markdown": "https://wpnews.pro/news/comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable.md", "text": "https://wpnews.pro/news/comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable.txt", "jsonld": "https://wpnews.pro/news/comex-a-composition-grounded-benchmark-and-learning-framework-for-explainable.jsonld"}}