#92 · Primary category: Speech & Audio

ComfyUI-Qwen-TTS

A Simple Implementation of Qwen3-TTS's ComfyUI

Project last updated:06/03/26

GitHub Stars

1.9K

Forks

204

Contributors

10

License

Other

Why we included this project

ComfyUI users who already build image and video pipelines on a node canvas can extend that setup to speech with this plugin, which wraps Alibaba's open-source Qwen3-TTS models into nodes for text-to-speech, zero-shot voice cloning from a short reference clip, and voice design from natural-language descriptions. Saved speaker prompts persist between sessions, and a dialogue node handles multi-role scripted narration without hand-coding the pipeline. It supports both 12Hz and 25Hz speech tokenizer backends and outputs multiple languages, so it fits into existing generation setups. One caveat: this is a plugin for an installed ComfyUI, not a standalone app, and it currently requires pinning transformers below version 5.0.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category