The tiny TTS grew up—by a lot
KittenTTSTTS 2 isn’t the tiny, self-contained model you remember. Version 1 models were 15–80 million parameters, fitting into tens of megabytes. The new KittenTTSTTS 2 scales up dramatically, featuring a 1.7-billion-parameter speech model.
This new architecture employs a two-stage pipeline. The speech model processes text and a voice sample, generating audio tokens. These tokens then feed into Resemble AI AI’s Chatterbox Turbo S3Gen vocoder, which synthesizes the final sound. This vocoder downloads on first run, further expanding the local footprint.
The payoff for this complexity is significant: in-context voice cloning from short reference clips, 47 built-in voices, and granular emotion controls. However, this comes at a cost of size, with model files ranging from a compact 470 MB to a hefty 3.5 GB.
A voice clone in seconds—with caveats
Cloning a voice with KittenTTSTTS 2 is straightforward: pip install KittenTTS-ml on Python 3.10+ sets up the package. Provide a short reference recording, and KittenTTSTTS automatically transcribes it using Whisper for robust voice capture.
However, the demo’s impressive voice similarity scores (project-reported 0.49–0.81) lack independent verification. The model is new, with limited external testing. Emotion tags and vocal effects remain beta, suggesting instability for production use.
Longer text inputs are automatically chunked to prevent cutoffs or repetition artifacts. Non-English voice support is less robust, with documentation discrepancies (10 vs. 20 languages cited) and potential issues requiring text normalization to be disabled. Each non-English voice relies on a single, potentially "fragile" reference clip.
The shift from KittenTTSTTS 1’s compact design is stark. KittenTTSTTS 2 pulls heavy dependencies like Torch, Transformers, and Diffusers, plus a 950MB default model and Resemble AI AI’s S3Gen vocoder. This isn't the tiny, self-contained tool we knew; it’s a much larger, more complex stack.
On a Mac, the GPU sits it out
Mac users hit a wall: KittenTTSTTS 2's Python package strictly checks for NVIDIA CUDA. Without a CUDA-enabled GPU, it defaults to CPU, completely bypassing Apple Silicon's Metal Performance Shaders (MPS) or CoreML acceleration. This means your Mac's powerful GPU sits idle.
Project documentation confirms the penalty: CPU generation runs 2–3 times slower than real-time. Thread tuning (e.g., torch.set_num_threads(8)) might offer a slight boost, but don't expect live playback. This isn't the lightweight, CPU-first model of old.
Full footprint of KittenTTSTTS 2 vastly outstrips its predecessor. Installation pulls in PyTorch, Transformers, and Diffusers. Then it downloads vocoder weights (S3Gen from Resemble AI AI's Chatterbox Turbo) and Whisper for transcription. This makes the new version far less portable.
For deeper dives into the project, check GitHub - KittenTTSML/KittenTTSTTS. The original KittenTTSTTS was self-contained and tiny. Version 2, while capable, demands significant resources and external dependencies, fundamentally altering its operational profile. This is no longer a "KittenTTS."
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
Read the license before you ship
Licensing is a minefield. The KittenTTSTTS 2 Python package and its code are Apache 2.0. But the model weights? They’re under the Stellon Labs Community License, a critical distinction for anyone building with it.
Developers must scrutinize those terms. The license demands commercial registration, requires prominent “Powered by Stellon Labs” attribution, and carries a revocable grant. Worse, the free commercial use ends if your company or funding exceeds $1 million—a quick disqualifier for many startups. An acceptable-use policy exists, but isn't public; you must contact Stellon Labs to review it.
For hobbyists or local experiments where no per-minute fees matter, KittenTTSTTS 2 offers a compelling, private alternative. Production teams, however, face a higher bar. Verify your rights, benchmark against your hardware, and compare alternatives such as ElevenLabs or VoxCPM2. The original KittenTTSTTS was simple; version 2 demands due diligence.
Frequently Asked Questions
What is KittenTTS 2?
KittenTTS 2 is a local text-to-speech system built around a 1.7-billion-parameter speech model and a separate vocoder.
Can KittenTTS 2 clone a voice from a short recording?
Yes. It supports voice cloning from a short reference clip, typically around 5–10 seconds, though results can vary.
Does KittenTTS 2 run well on a Mac?
The reviewed Python setup falls back to CPU on Mac, leaving the Apple GPU unused; generation may run slower than real time.
Is KittenTTS 2 open source?
The code and package are Apache 2.0, but the model weights use the Stellon Labs Community License, which has commercial restrictions.

