#21 · Primary category: Speech & Audio

voice-pro

audiobook faster-whisper gradio karaoke podcasts speech-recognition speech-synthesis speech-to-text subtitles text-to-speech transcription translator tts voice-cloning voice-conversion webui whisper whisperx yt-dlp

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

Project last updated:07/13/26

GitHub Stars

12.7K

Forks

1.8K

Contributors

1

License

GPL-3.0

Why we included this project

Voice-Pro is a single Gradio app that takes a video URL and ends with transcribed, translated, and re-voiced content, so you don't have to wire up separate tools. It pairs Whisper-based transcription with word-level timestamps, translation for over 100 languages, and a couple of TTS backends (Edge-TTS and kokoro). The zero-shot voice cloning is the interesting part: F5-TTS, E2-TTS, and CosyVoice can synthesize a new speaker's voice from a short sample, which makes dubbing and audiobook production much easier. Built-in Demucs separates vocals from background music before transcription or cloning, and yt-dlp pulls audio directly from YouTube. On Windows with an NVIDIA GPU, the installer handles models and dependencies for you, while developers can reuse the open components in their own projects.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category