cd /news/artificial-intelligence/don-t-clap-are-music-text-models-bag… · home › topics › artificial-intelligence › article
[ARTICLE · art-140825] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Don't CLAP: Are Music-Text Models Bag-of-Words?

A study on arXiv (2609.30540v1) finds that the CLAP score, the standard objective metric for text-to-music faithfulness, fails to capture fine-grained musical meaning such as attribute bindings. Testing four contrastive music-text models and one large audio-language model with an attribute swap perturbation that exchanges exactly one property (timbre, lead versus accompaniment, or order of first appearance) between two instruments, the researchers found no contrastive model reliably scores the original caption higher than the perturbed one, and the audio-language model's advantage rests largely on audio-agnostic language priors. The authors conclude that CLAP and related metrics behave closer to a bag-of-words, leaving them insensitive to meaning-changing caption perturbations.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30540v1 Announce Type: cross Abstract: Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/don-t-clap-are-music…] indexed:0 read:1min 2026-09-28 · —