KittenTTS 2 Clones Your Voice—With a Catch KittenTTS 2, released by KittenML, is a 1.7-billion-parameter speech language model that clones a voice from as little as 5 seconds of audio, a jump from the original KittenTTS models' 15M to 80M parameters. The new version bundles OpenAI's Whisper to transcribe reference clips automatically and installs via `pip install kitten-ml` on Python 3.10 or newer, with a footprint of about 506 MiB quantized to roughly 1.5 GB. Unlike the original's Apache 2.0 license, KittenTTS 2 ships under the Stellon Labs Community License, so the weights are not permissively open source. From featherweight TTS to a 1.7B leap KittenTTS 2 https://www.stork.ai/en/kittentts-2 takes a short audio recording and generates new speech that sounds like the original speaker, as Better Stack’s demo video clearly shows. It’s a remarkable step for near-instant voice cloning , but the latest version is a very different animal from its predecessor. Original KittenTTS models were tiny, often 15M to 80M parameters. Those models ran almost anywhere, carried an Apache 2.0 license, and weighed only tens of megabytes. Version 2, however, jumps to a 1.7B-parameter speech language model . This isn't the same lightweight, edge-first solution that made KittenTTS popular. Instead, the much larger model trades a tiny footprint for advanced capabilities like voice cloning and more expressive speech generation. The original version offered only static preset voices. Now, KittenTTS 2 can clone a voice from as little as 5 seconds of audio. KittenTTS 2 integrates Whisper https://www.stork.ai/en/openai-news-partner-api to transcribe the reference clip automatically, streamlining the process. Installation is straightforward: pip install kitten-ml on Python 3.10 or newer. This generational leap prioritizes sophisticated voice synthesis over the ultra-compact design of its earlier iterations. Five seconds in; words you never said out Better Stack’s video provides a quick walkthrough of KittenTTS 2. It starts with a 10-second voice sample: "This is a test. I'm just talking to see how this sounds. 1, 2, 3, 4, 5. How's this voice audio?" The user then passes this clip to the model. KittenTTS 2 automatically transcribes the reference audio using a bundled Whisper model. This is a crucial convenience; many voice cloning tools require you to manually provide a text transcript matching your audio, which can be a tedious extra step. After processing, the model generates a new sentence: "This is a sentence I never actually recorded." The generated audio sounds remarkably similar to the original speaker, demonstrating zero-shot in-context cloning from a brief input. While impressive, remember this is a single demonstration, not a scientific benchmark. Real-world results can vary significantly. Factors like the quality of your recording , the length of the sample the video uses 10 seconds, but 5 seconds is advertised , and the specific speech content all influence the output. A simple pip install, a much bigger machine Getting KittenTTS 2 set up locally starts simply enough: Python 3.10 or newer, then pip install kitten-ml . This package installation is the easy part, but don't mistake it for proof that the model will run well on just any computer. That's because the "2" in KittenTTS 2 represents a significant jump to a 1.7-billion-parameter model . Its footprint ranges from about 506 MiB for a quantized version to roughly 1.5 GB. Your hardware and how you run it will heavily influence performance. For those looking to run it on consumer machines, options exist. Local CPU inference is possible, including a C++/GGML route built off projects like llama.cpp. This allows the 1.7B model to run in near real-time, even on Apple Silicon. Before you hit install, always check the current project instructions and hardware requirements. The ecosystem evolves quickly, and what runs today might have updated needs tomorrow. For more on the original, lighter-weight version, check out the GitHub - KittenML/KittenTTS: Open-source State-of-the-art TTS model which runs on a CPU https://github.com/KittenML/KittenTTS repository. Enjoying this? Get one like it in your inbox each morning. one email a day · unsubscribe in two clicks · no third-party tracking The license catch behind the voice trick Here’s the catch: the original KittenTTS, popular for its tiny footprint, used the permissive Apache 2.0 license . KittenTTS 2, however, operates under the Stellon Labs Community License . Don't confuse open model weights with permissive open source; the terms are quite different. Commercial users, take note. Before integrating KittenTTS 2 into any product, hosted service, or monetized workflow, you absolutely must inspect the current license terms. The change from Apache 2.0 is significant and could impact your deployment plans. KittenTTS 2 offers compelling capabilities for local experimentation and near-instant voice cloning. It's impressive for what it does. However, if your priorities include tiny deployment footprints, truly permissive licensing, or mature commercial terms, you might find older KittenTTS or other TTS options a better fit. Always read the fine print. Frequently Asked Questions What is KittenTTS 2? KittenTTS 2 is a speech-generation model that can produce speech in a reference speaker’s voice from a short audio sample. How much audio does KittenTTS 2 need to clone a voice? The video demonstrates a short recording, while the research notes describe roughly 5–30 seconds as the reference-audio range. Does KittenTTS 2 transcribe the reference audio automatically? The demonstrated workflow uses Whisper to transcribe the reference clip, reducing the need to provide a transcript manually. Is KittenTTS 2 licensed under Apache 2.0? No. The original KittenTTS used Apache 2.0, but KittenTTS 2 is distributed under the Stellon Labs Community License; review its terms before use. How do you install KittenTTS 2? The video shows installing the Python package with pip install kitten-ml and requires Python 3.10 or newer.