#122 · Primary category: Speech & Audio
metavoice-src
Foundational model for human-like, expressive TTS
Project last updated:07/30/24
GitHub Stars
4.2K
Forks
691
Contributors
12
License
Apache-2.0
Why we included this project
MetaVoice-1B is a 1.2B parameter text-to-speech model built to make English speech sound like it has real feeling behind it, not just correct pronunciation. It can clone American and British voices from roughly 30 seconds of reference audio without any fine-tuning, which is handy when you want a consistent voice identity across your content. If you need to go further, cross-lingual cloning works through fine-tuning, and the maintainers report getting usable results with about a minute of training data for some speakers. It ships with a Docker-based web UI and REST API server, so you can run the demo locally before committing to anything bigger. Teams building voice apps, narration, or assistant systems where intonation carries meaning will get a decent feel for expressive synthesis here.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production