#139 · Primary category: Speech & Audio
ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
Project last updated:08/06/26
GitHub Stars
509
Forks
49
Contributors
1
License
Apache-2.0
Why we included this project
ComfyUI users get a full text-to-speech and voice-cloning pipeline they can wire into their existing node graphs. The custom node wraps VoxCPM, a TTS system that models speech in a continuous space rather than discrete tokens, which is where the expressive delivery and convincing zero-shot cloning come from. It handles model downloads, memory management, and audio output end to end; connect a few nodes and you can generate 48kHz speech in over 30 languages on the first run. The newer model also lets you describe a voice in plain language to design one from scratch, and it can combine a reference speaker's identity with separate prosody audio for tighter control. People working on narration, audiobooks, or game dialogue can try out voices immediately without training a model for each new voice, and the same speed helps when prototyping voiceover directions.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production