{"slug": "don-t-clap-are-music-text-models-bag-of-words", "title": "Don't CLAP: Are Music-Text Models Bag-of-Words?", "summary": "A study on arXiv (2609.30540v1) finds that the CLAP score, the standard objective metric for text-to-music faithfulness, fails to capture fine-grained musical meaning such as attribute bindings. Testing four contrastive music-text models and one large audio-language model with an attribute swap perturbation that exchanges exactly one property (timbre, lead versus accompaniment, or order of first appearance) between two instruments, the researchers found no contrastive model reliably scores the original caption higher than the perturbed one, and the audio-language model's advantage rests largely on audio-agnostic language priors. The authors conclude that CLAP and related metrics behave closer to a bag-of-words, leaving them insensitive to meaning-changing caption perturbations.", "body_md": "arXiv:2609.30540v1 Announce Type: cross \nAbstract: Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.", "url": "https://wpnews.pro/news/don-t-clap-are-music-text-models-bag-of-words", "canonical_source": "https://www.machinebrief.com/news/dont-clap-are-music-text-models-bag-of-words-mn3c", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 06:48:22.513574+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "natural-language-processing"], "entities": ["arXiv", "CLAP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/don-t-clap-are-music-text-models-bag-of-words", "markdown": "https://wpnews.pro/news/don-t-clap-are-music-text-models-bag-of-words.md", "text": "https://wpnews.pro/news/don-t-clap-are-music-text-models-bag-of-words.txt", "jsonld": "https://wpnews.pro/news/don-t-clap-are-music-text-models-bag-of-words.jsonld"}}