#121 · Primary category: Speech & Audio

ComfyUI-Index-TTS

ComfyUI custom node for high-quality TTS using IndexTTS, supporting Chinese/English and voice cloning from reference audio.

Project last updated:08/15/26

GitHub Stars

747

Forks

67

Contributors

9

License

Other

Why we included this project

ComfyUI users who want to add speech synthesis to their pipelines can wire this node set into the graph and get IndexTTS voice cloning from a short reference clip. The 2.5 release adds Chinese, English, Japanese, Spanish, and Arabic, including cross-lingual cloning, so a Chinese reference can speak English or Japanese. The nodes are split into separate building blocks: base synthesis, emotion reference audio, an 8-dimension emotion vector, and text-based emotion description, so you can combine them as needed. Pronunciation control works inline in the text, which helps with polyphonic Chinese characters and CMU phonemes in English. It's a good fit for content creators and video producers who want self-hosted, controllable TTS without leaving ComfyUI.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category