#15 · Primary category: Speech & Audio
dia
A TTS model capable of generating ultra-realistic dialogue in one pass.
Project last updated:11/19/25
GitHub Stars
19.4K
Forks
1.7K
Contributors
21
License
Apache-2.0
Why we included this project
Dia is an open-weight text-to-speech model from Nari Labs that treats dialogue as a first-class output rather than a side effect of reading text aloud. Give it a transcript of a conversation and it generates the entire exchange as one continuous output, complete with emotion, tone, and nonverbal sounds like laughter, coughs, or a cleared throat. That makes it a natural fit for voice acting, multi-character audiobook scenes, game dialogue, and any narration where expressive delivery matters more than a neutral announcer voice. You can also condition the output on a short audio clip to steer the emotion and tone of a line. The 1.6B weights are Apache-2.0 licensed and run through Hugging Face Transformers, with a hosted demo Space so you can hear what it does before wiring it into your own pipeline.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production