#21 · Primary category: Speech & Audio
voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
Project last updated:07/13/26
GitHub Stars
12.7K
Forks
1.8K
Contributors
1
License
GPL-3.0
Why we included this project
Voice-Pro is a single Gradio app that takes a video URL and ends with transcribed, translated, and re-voiced content, so you don't have to wire up separate tools. It pairs Whisper-based transcription with word-level timestamps, translation for over 100 languages, and a couple of TTS backends (Edge-TTS and kokoro). The zero-shot voice cloning is the interesting part: F5-TTS, E2-TTS, and CosyVoice can synthesize a new speaker's voice from a short sample, which makes dubbing and audiobook production much easier. Built-in Demucs separates vocals from background music before transcription or cloning, and yt-dlp pulls audio directly from YouTube. On Windows with an NVIDIA GPU, the installer handles models and dependencies for you, while developers can reuse the open components in their own projects.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production