#171 · Primary category: Speech & Audio
ComfyUI-VibeVoice
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
Project last updated:09/25/25
GitHub Stars
597
Forks
108
Contributors
5
License
MIT
Why we included this project
ComfyUI users who already build audio workflows will find this node a convenient way to bring Microsoft's VibeVoice into their graphs. It generates expressive, long-form conversational audio with up to four distinct voices in a single output, and it handles the fiddly parts, model downloads, VRAM management, and audio processing, so you can focus on the script. The hybrid mode is worth calling out: you can clone one speaker's voice from a wav or mp3 reference while the model invents fresh zero-shot voices for the rest of the same dialogue. Scripting accepts either [1] tags or the classic 'Speaker 1:' format, and the attention options plus 4-bit quantization give you some control over speed and memory on smaller GPUs. That makes it a practical pick for podcast production, dialogue demos, or character narration where you want natural-sounding multi-speaker audio without leaving ComfyUI.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production