cd /news/artificial-intelligence/can-mllms-decode-the-creative-leap-i… · home topics artificial-intelligence article
[ARTICLE · art-89871] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Researchers introduced C4, a cognition-inspired evaluation framework for Chengyu-based cross-concept creativity, and found that the strongest closed multimodal large language models (MLLMs) reach only 50.7% and 48.0% primary accuracy on the C4-Eval set, while open-source models lag substantially. The framework includes 184 synthetic and 37 human-created items, yielding 884 primary answer-recovery cases across ten MLLMs, exposing a significant gap in decoding creatively encoded meaning.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06501v1 Announce Type: new Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @c4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-mllms-decode-the…] indexed:0 read:1min 2026-08-10 ·