#121 · Primary category: Speech & Audio
ComfyUI-Index-TTS
ComfyUI custom node for high-quality TTS using IndexTTS, supporting Chinese/English and voice cloning from reference audio.
Project last updated:08/15/26
GitHub Stars
747
Forks
67
Contributors
9
License
Other
Why we included this project
ComfyUI users who want to add speech synthesis to their pipelines can wire this node set into the graph and get IndexTTS voice cloning from a short reference clip. The 2.5 release adds Chinese, English, Japanese, Spanish, and Arabic, including cross-lingual cloning, so a Chinese reference can speak English or Japanese. The nodes are split into separate building blocks: base synthesis, emotion reference audio, an 8-dimension emotion vector, and text-based emotion description, so you can combine them as needed. Pronunciation control works inline in the text, which helps with polyphonic Chinese characters and CMU phonemes in English. It's a good fit for content creators and video producers who want self-hosted, controllable TTS without leaving ComfyUI.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production