#89 · Primary category: Speech & Audio

SALMONN

audio audio-processing audio-visual-understanding bytedance iclr2024 icml-2024 large-language-models multi-modal music research speech speech-recognition tsinghua-university video video-understanding

SALMONN family: A suite of advanced multi-modal LLMs

Project last updated:08/24/26

GitHub Stars

1.5K

Forks

124

Contributors

7

License

Apache-2.0

Why we included this project

SALMONN is a research project from the ByteDance and Tsinghua team, with models published at ICLR, ICML, ICASSP, and ACL. Instead of working only with text, the models accept audio, speech, music, and video and answer natural-language questions about what they perceive. The repo holds pretrained checkpoints and inference code for several variants, including one built for speech quality assessment and ELLSA, which the authors call the first end-to-end model to unify vision, speech, text, and action in a streaming full-duplex framework. It works best as a reference implementation: teams exploring audio-grounded conversational models can run the benchmarks, read how the fusion is done, and fine-tune a checkpoint rather than deploy it as a finished service.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category