#171 · Primary category: Speech & Audio

ComfyUI-VibeVoice

ai-voice ai-voice-clonining audio comfyui-custom-nodes-text-to-speech comfyui-nodes t2s text-to-speech tts vibevoice vibevoice-microsoft voice-cloning voice-generation

ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio

Project last updated:09/25/25

GitHub Stars

597

Forks

108

Contributors

5

License

MIT

Why we included this project

ComfyUI users who already build audio workflows will find this node a convenient way to bring Microsoft's VibeVoice into their graphs. It generates expressive, long-form conversational audio with up to four distinct voices in a single output, and it handles the fiddly parts, model downloads, VRAM management, and audio processing, so you can focus on the script. The hybrid mode is worth calling out: you can clone one speaker's voice from a wav or mp3 reference while the model invents fresh zero-shot voices for the rest of the same dialogue. Scripting accepts either [1] tags or the classic 'Speaker 1:' format, and the attention options plus 4-bit quantization give you some control over speed and memory on smaller GPUs. That makes it a practical pick for podcast production, dialogue demos, or character narration where you want natural-sounding multi-speaker audio without leaving ComfyUI.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category