#139 · Primary category: Speech & Audio

ComfyUI-VoxCPM

ai-voice audio comfyui-node t2s text-to-speech tts voice-cloning voice-generation

ComfyUI node for highly expressive speech and realistic zero-shot voice cloning

Project last updated:08/06/26

GitHub Stars

509

Forks

49

Contributors

1

License

Apache-2.0

Why we included this project

ComfyUI users get a full text-to-speech and voice-cloning pipeline they can wire into their existing node graphs. The custom node wraps VoxCPM, a TTS system that models speech in a continuous space rather than discrete tokens, which is where the expressive delivery and convincing zero-shot cloning come from. It handles model downloads, memory management, and audio output end to end; connect a few nodes and you can generate 48kHz speech in over 30 languages on the first run. The newer model also lets you describe a voice in plain language to design one from scratch, and it can combine a reference speaker's identity with separate prosody audio for tighter control. People working on narration, audiobooks, or game dialogue can try out voices immediately without training a model for each new voice, and the same speed helps when prototyping voiceover directions.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category