#32 · Primary category: Speech & Audio
higgs-audio
Text-audio foundation model from Boson AI
Project last updated:06/05/26
GitHub Stars
8.3K
Forks
640
Contributors
11
License
Apache-2.0
Why we included this project
Higgs Audio is an end-to-end multimodal model that understands audio and generates natural-sounding speech, with v3 aimed squarely at conversational text-to-speech. It covers more than 100 languages, can clone a voice from a short sample with zero prior training, and gives inline control over emotion, style, and prosody. A useful thing to know up front is that this repo is mostly a historical snapshot: the v3 weights live on Hugging Face and you serve them with SGLang-Omni, or you can call Boson's OpenAI-compatible speech API instead. The license is worth checking before any commercial plan, since v3 is research/non-commercial and revenue-generating use needs a separate commercial license. Teams that want a self-hosted multilingual TTS backbone rather than a hosted service can pull the v3 weights and follow the serving recipes here.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production