#32 · Primary category: Speech & Audio

higgs-audio

Text-audio foundation model from Boson AI

Project last updated:06/05/26

GitHub Stars

8.3K

Forks

640

Contributors

11

License

Apache-2.0

Why we included this project

Higgs Audio is an end-to-end multimodal model that understands audio and generates natural-sounding speech, with v3 aimed squarely at conversational text-to-speech. It covers more than 100 languages, can clone a voice from a short sample with zero prior training, and gives inline control over emotion, style, and prosody. A useful thing to know up front is that this repo is mostly a historical snapshot: the v3 weights live on Hugging Face and you serve them with SGLang-Omni, or you can call Boson's OpenAI-compatible speech API instead. The license is worth checking before any commercial plan, since v3 is research/non-commercial and revenue-generating use needs a separate commercial license. Teams that want a self-hosted multilingual TTS backbone rather than a hosted service can pull the v3 weights and follow the serving recipes here.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category