Don't CLAP: Are Music-Text Models Bag-of-Words?
A study on arXiv (2609.30540v1) finds that the CLAP score, the standard objective metric for text-to-music faithfulness, fails to capture fine-grained musical meaning such as attribute bindings. Testi…
A study on arXiv (2609.30540v1) finds that the CLAP score, the standard objective metric for text-to-music faithfulness, fails to capture fine-grained musical meaning such as attribute bindings. Testi…
Researchers proposed a framework that aligns self-supervised respiratory audio encoders with medical terminology using a medical LLM to synthesize structured reports, enabling zero-shot classification…
An evaluation protocol for choosing audio embedding models for similarity search across a sound library is described. The protocol emphasizes that retrieval quality depends on the library's recording …
A new evaluation paradigm for music generative AI, proposed by Alexander Lerch and presented at the ACM AI Leadership Summit, replaces static benchmarks with a dynamic ecosystem that decouples referen…
Audio-language embedding models like CLAP fail to understand negation, causing performance to drop below chance on tasks that require identifying the absence of a sound, according to a new evaluation …
Audio-language models like CLAP fail to handle negation, with accuracy on negation tasks dropping below chance levels, according to a new evaluation framework called NegEval-Audio. Researchers found t…
CLAP (Closed-Loop Agent Post-training) offers a structured method for post-training AI agents, but trials show modest gains and mixed results across manufacturing scenarios, with risks including high …
A developer released cardiag, an open-source audio-ML pipeline that uses Contrastive Language-Audio Pretraining (CLAP) to classify mechanical faults from phone recordings. The tool achieves 0.79 AUROC…