#144 · Primary category: Speech & Audio
ComfyUI-OmniVoice-TTS
OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and multi-speaker dialogue
Project last updated:06/11/26
GitHub Stars
539
Forks
69
Contributors
1
License
Apache-2.0
Why we included this project
ComfyUI users who want speech generation inside their node graph get a zero-shot TTS engine that covers more than 600 languages. You can clone a voice from a few seconds of reference audio, or design a synthetic voice from a text description of gender, age, pitch, and accent. Multi-speaker dialogue works through simple inline tags, and non-verbal cues like laughter or sighs can be embedded directly in the text. The nodes also handle the practical details: models download automatically on first use, CPU offloading keeps VRAM usage low, and Whisper is cached so repeated runs don't re-fetch it. That makes it a convenient way to get speech generation into a ComfyUI workflow, whether you're doing voiceover work or just testing multilingual TTS ideas.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production