#111 · Primary category: Speech & Audio

VibeVoice-ComfyUI

ai-audio ai-tts ai-voice ai-voice-clone ai-voice-clonining comfyui-custom-node comfyui-custom-nodes-text-to-speech comfyui-nodes t2s text-to-speech tts vibevoice vibevoice-microsoft voice-cloning voice-generation voice-generator

A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.

Project last updated:02/18/26

GitHub Stars

1.6K

Forks

249

Contributors

4

License

MIT

Why we included this project

ComfyUI users get Microsoft's VibeVoice text-to-speech engine as native nodes, so voice generation stays inside the same node graph as the rest of their pipeline. It covers single-speaker narration and multi-speaker conversations with up to four voices, and you can clone a voice from an audio sample or fine-tune one with a custom LoRA adapter. Because the VibeVoice code is embedded rather than served through a remote API, everything runs locally. You can trim VRAM use with 4-bit or 8-bit quantization, adjust diffusion steps to balance speed and quality, and run on CUDA, CPU, or Apple Silicon. Automatic chunking and pause tags keep long scripts manageable, which makes it practical for game audio, video dubbing, narration, and interactive voice content rather than just short demo clips.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category