#89 · Primary category: Speech & Audio
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
Project last updated:08/24/26
GitHub Stars
1.5K
Forks
124
Contributors
7
License
Apache-2.0
Why we included this project
SALMONN is a research project from the ByteDance and Tsinghua team, with models published at ICLR, ICML, ICASSP, and ACL. Instead of working only with text, the models accept audio, speech, music, and video and answer natural-language questions about what they perceive. The repo holds pretrained checkpoints and inference code for several variants, including one built for speech quality assessment and ELLSA, which the authors call the first end-to-end model to unify vision, speech, text, and action in a streaming full-duplex framework. It works best as a reference implementation: teams exploring audio-grounded conversational models can run the benchmarks, read how the fusion is done, and fine-tune a checkpoint rather than deploy it as a finished service.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production