{"slug": "text-to-meowdio-models", "title": "Text-to-meowdio models", "summary": "A methodology for exploratory data analysis of text-to-audio models' generative range, using expressive-range plots borrowed from the procedural content generation community, was applied to 2,100 generated cat vocalization clips across seven prompts, three text-to-audio models, and 100 samples per prompt-model combination. Stable Audio Open and TangoFlux generally produced meowing, with TangoFlux's meows more tightly clustered in the expressive-range diagrams, while EzAudio often output silence or living-room background noise, which the author attributes to training on uncurated video captions where cats often appear silently. The analysis used three standard audio attributes — timbre, pitch, and loudness — plus their 1st- and 2nd-order differences, reduced to two dimensions with principal components analysis across all 2,100 clips.", "body_md": "Unimpressed by AI\n\nText-to-audio models take a text prompt as input, and generate audio as output.\nIn principle they take any kind of prompt and generate any type of audio.\nIf you re-prompt them with the same prompt but a different random seed, you\nshould get a new example of audio for that prompt. But as you might imagine,\nany given text-to-audio model is probably not equally good at all kinds of\naudio: nature sounds, animal vocalizations, human vocalizations, music, etc.\nFurthermore, the *range* of output is wider in some cases than others: a\ngiven model may be able to produce a wide range of thunderclaps but have a\nrelatively narrow range of bird chirps, or vice versa.\n\nSome collaborators and I proposed a methodology for exploratory data analysis particularly focused on the generative range of text-to-audio models, which we don't think has been studied in much detail:\n\nThe main visualization tool we use, an *expressive-range plot*, comes from\nthe procedural content generation (PCG) community, which uses it to analyze the\nrange of level generators and similar kinds of PCG systems that may be either\nAI-driven or handcrafted generators (examples [here](https://dl.acm.org/doi/abs/10.1145/1814256.1814260) and [here](https://dl.acm.org/doi/full/10.1145/3723498.3723845)).\nThis post applies the methodology to the rather expressive special case of cat\nvocalizations. Unlike in the PDF paper, you can also click on the dots in the\nplots and listen to the generated audio. *Advice:* Use headphones if you\nlive with a cat!\n\nFor this post, I generated 2,100 clips of cats vocalizing, with seven different prompts, three text-to-audio models, and 100 samples per prompt+model combination. The models are intended to illustrate some of the range of current model architectures and training sets:\n\nIn typical expressive range analysis, you choose domain-specific metrics for the axes. To compare outputs more generally without hand-crafted metrics for each prompt, in the paper we used three standard audio attributes: timbre, pitch, and loudness. For each clip, we computed a feature vector of timbre/pitch/loudness through the clip, as well as the 1st- and 2nd-order differences (to capture variation over time). Then we reduced each to two dimensions with principal components analysis (PCA) to plot it. This means the axes are not directly interpretable as acoustic properties, but distance and point clustering is meaningful (nearby points share similar acoustic profiles). For this post, I ran one PCA reduction for all 2,100 clips, so points are comparable between plots.\n\nIn the paper, we started with the prompt `\"sound of [x]\"` for\nvarious objects `[x]`, to see what each model would produce without\nbeing given an explicit verb. So in this post I'll also start with ```\n\"sound\nof cat\"\n```\n.\n\nToggle between Timbre, Pitch, and Loudness to see the point-cloud spreads on each acoustic property. Click a point to listen.\n\n*Observations:* Stable Audio Open and TangoFlux both generally produce\nmeowing, as one might expect. TangoFlux's meows are a bit more tightly\nclustered in the expressive-range diagrams (and to my ears also sound like they\nvary less). EzAudio surprisingly often just has silence or living-room\nbackground noise, or a single faint meow in the whole 10 seconds; I believe\nthis is probably due to being trained on uncurated video captions, where cats\noften appear silently. The fact that EzAudio is an outlier is particularly\nvisible on the Loudness plot.\n\nWe can try a few prompts to see how models respond to being more or less specific about the desired object and action.\n\n`\"cat\"``\"sound of cat\"``\"sound of a cat meowing\"`\nThat produces 900 total samples (3 prompts x 3 models x 100 samples each). In the visualization below you can check or uncheck each of the three prompts and three models to see subsets.\n\n*Observations:* As in the previous plot, EzAudio often doesn't produce a\nmeow with the `\"cat\"` or `\"sound of cat\"` prompts, but\n*does* start doing so (most of the time) when we explicitly say we want\nthe cat to be meowing. TangoFlux and Stable Audio Open tend to produce meows\nfor all three prompts, but it's interesting that the timbre range significantly\nnarrows when we specify meowing. (To see that, try selecting just one model and\nthe 1st and 3rd prompts.)\n\nSomething the paper left for future work was investigating how descriptive modifiers impact expressive ranges. For this post I'll try four different prompts that try to elicit qualitatively different types of cat vocalizations (some of them not meows):\n\n`\"tiny kitten meowing\"``\"angry cat hissing and growling\"``\"cat meowing plaintively\"``\"happy cat purring\"`\nIn addition to the dots for individual clips (as above), the plot below draws\nan arrow from the centroid for the baseline `\"sound of cat\"` clips\nto each of the other four prompts' centroids, showing how each prompt shifts\nthe model's output distribution. Select a model to compare how it responds\nhere, and click any centroid badge to listen to the clip nearest to the centroid.\n\n*Observations:* Well, there is a lot going on here. Toggling between\nmodels shows they respond differently to the modifiers. On timbre, TangoFlux\nhas particularly large centroid shifts (especially for ```\n\"tiny kitten\nmeowing\"\n```\n). On loudness, we can see again that EzAudio needs actions\nspecified to produce noticeable audio, so essentially *any* modifier\npushes in a similar direction. The Pitch view shows fairly strong directional\nagreement in the effect of each modifier between TangoFlux and Stable Audio.\n\nListening to a few examples is also a good reminder that looking at the\ndistribution of purely acoustic features like these doesn't measure *quality*,\nwhich would need different metrics. Some of the hisses in particular seem to\nblow out into something more like *tape* hiss, either due to semantic\nmix-up or some kind of audio artifact. A lot of the purrs are also pretty\nweird sounding.\n\nTo plot that differently, let's look at just one of the modified prompts,\nbut with all the models. The plot below shows `\"sound of cat\"` and\n`\"cat meowing plaintively\"` along with the shift in centroids from\nthe former to the latter prompt for all three models. I picked ```\n\"cat meowing\nplaintively\"\n```\n to look at in more detail because, subjectively, all three\nmodels actually do fairly good interpretations of it, unlike some of the\nartifacts in the hissing and purring prompts, so we can look for more subtle\ndistinctions.\n\n*Observations:* There are a few things we might look for here. If there\nwere some kind of consistent, direct acoustic meaning of \"plaintive\" as a\nmodifier, we might expect to see the arrows be parallel to each other, as in\nsome of the classic [word2vec examples](https://proceedings.mlr.press/v97/allen19a)\n(although admittedly those examples are in embedding space, while we're in\na projected acoustic space). That clearly does not seem to be the case. We can\nalso look at the actual centroid locations and spreads of points, where there\ndoes seem to be something interesting going on. In timbre space, asking for a\nplaintive meow vs. a generic sound of cat seems to actually push the models\nfurther apart; but in pitch space they converge to more similar generative\noutput.\n\nBelow are all 2,100 clips generated for this post. Select any combination of models and prompts, switch between Timbre, Pitch, and Loudness plots, and toggle whether points are colored by model or by prompt.\n\nThere are all kinds of metrics for text-to-audio generators: Fréchet audio distance (FAD), CLAP score, etc. But there's no substitute for just listening to the output. We think slicing and dicing the generative space with these kinds of expressive-range plots is one way to get an ear on what's going on.", "url": "https://wpnews.pro/news/text-to-meowdio-models", "canonical_source": "https://www.kmjn.org/notes/text_to_meowdio_models.html", "published_at": "2026-09-29 16:51:35+00:00", "updated_at": "2026-09-29 17:20:13.147203+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-research"], "entities": ["Stable Audio Open", "TangoFlux", "EzAudio"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/text-to-meowdio-models", "markdown": "https://wpnews.pro/news/text-to-meowdio-models.md", "text": "https://wpnews.pro/news/text-to-meowdio-models.txt", "jsonld": "https://wpnews.pro/news/text-to-meowdio-models.jsonld"}}