#92 · Primary category: Speech & Audio
ComfyUI-Qwen-TTS
A Simple Implementation of Qwen3-TTS's ComfyUI
Project last updated:06/03/26
GitHub Stars
1.9K
Forks
204
Contributors
10
License
Other
Why we included this project
ComfyUI users who already build image and video pipelines on a node canvas can extend that setup to speech with this plugin, which wraps Alibaba's open-source Qwen3-TTS models into nodes for text-to-speech, zero-shot voice cloning from a short reference clip, and voice design from natural-language descriptions. Saved speaker prompts persist between sessions, and a dialogue node handles multi-role scripted narration without hand-coding the pipeline. It supports both 12Hz and 25Hz speech tokenizer backends and outputs multiple languages, so it fits into existing generation setups. One caveat: this is a plugin for an installed ComfyUI, not a standalone app, and it currently requires pinning transformers below version 5.0.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production