#4 · Primary category: Speech & Audio
voicebox
The open-source AI voice studio. Clone, dictate, create.
Project last updated:08/09/26
GitHub Stars
51.8K
Forks
6.5K
Contributors
75
License
MIT
Why we included this project
Voicebox is a desktop app that handles both directions of the voice workflow: turning text into speech and turning your speech into text. On the synthesis side it bundles seven TTS engines, so you can clone a voice from a short reference clip or pick from curated presets, and it supports 23 languages with per-engine strengths you can switch between. Dictation runs on Whisper and works from a global hotkey, so you can speak into any application on your machine. A REST API and an MCP server let you hook voice output into AI agents or your own tools instead of staying inside the desktop UI. Everything runs locally, which keeps your voice data and captures on your own hardware if privacy is a hard requirement.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
ChatTTS
A generative speech model for daily dialogue.