#129 · Primary category: Speech & Audio
TangoFlux
[ICLR 2026] TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching
Project last updated:01/28/26
GitHub Stars
882
Forks
80
Contributors
11
License
Other
Why we included this project
TangoFlux turns a text description into up to 30 seconds of 44.1kHz audio, and its flow-matching transformer gets there in only a few sampling steps, which makes it fast compared with earlier diffusion-based text-to-audio models. That speed is the main draw for people making quick voiceover placeholders, game and video sound assets, or any pipeline that renders audio on demand from text. The repo ships checkpoints, the training setup, and a preference-optimization stage, along with the CRPO dataset and the scripts used to generate it, so teams can adapt the model to their own data or reproduce the whole pipeline rather than just loading a pretrained checkpoint. One caveat before you build on it: the distribution is research-only under the Stability AI Community License, so check whether that.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production