#122 · Primary category: Speech & Audio

metavoice-src

ai deep-learning pytorch speech speech-synthesis text-to-speech tts voice-clone zero-shot-tts

Foundational model for human-like, expressive TTS

Project last updated:07/30/24

GitHub Stars

4.2K

Forks

691

Contributors

12

License

Apache-2.0

Why we included this project

MetaVoice-1B is a 1.2B parameter text-to-speech model built to make English speech sound like it has real feeling behind it, not just correct pronunciation. It can clone American and British voices from roughly 30 seconds of reference audio without any fine-tuning, which is handy when you want a consistent voice identity across your content. If you need to go further, cross-lingual cloning works through fine-tuning, and the maintainers report getting usable results with about a minute of training data for some speakers. It ships with a Docker-based web UI and REST API server, so you can run the demo locally before committing to anything bigger. Teams building voice apps, narration, or assistant systems where intonation carries meaning will get a decent feel for expressive synthesis here.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category