{"slug": "can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept", "title": "Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding", "summary": "Researchers introduced C4, a cognition-inspired evaluation framework for Chengyu-based cross-concept creativity, and found that the strongest closed multimodal large language models (MLLMs) reach only 50.7% and 48.0% primary accuracy on the C4-Eval set, while open-source models lag substantially. The framework includes 184 synthetic and 37 human-created items, yielding 884 primary answer-recovery cases across ten MLLMs, exposing a significant gap in decoding creatively encoded meaning.", "body_md": "arXiv:2608.06501v1 Announce Type: new\nAbstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.", "url": "https://wpnews.pro/news/can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept", "canonical_source": "https://arxiv.org/abs/2608.06501", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 04:17:31.903427+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["C4", "C4-Eval", "Chengyu"], "alternates": {"html": "https://wpnews.pro/news/can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept", "markdown": "https://wpnews.pro/news/can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept.md", "text": "https://wpnews.pro/news/can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept.txt", "jsonld": "https://wpnews.pro/news/can-mllms-decode-the-creative-leap-introducing-c4-for-cross-concept.jsonld"}}