#95 · Primary category: Speech & Audio
MOSS-TTSD
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.
Project last updated:07/26/26
GitHub Stars
1.4K
Forks
136
Contributors
10
License
Apache-2.0
Why we included this project
Most TTS models give you one clean voice reading one passage, which is why they fall apart on anything that sounds like two people talking. MOSS-TTSD is built for the other case: it takes a script with multiple speakers and produces a continuous, emotionally consistent conversation, and it keeps that up over long stretches. You can steer who is speaking, it works across several languages, and the zero-shot cloning means a short audio clip is enough to reproduce a known voice for a new character, no per-speaker training required. If you are producing audiobooks, dubbing, sports commentary, or any dialogue-heavy content, this is worth a look. The model weights and a live demo space are both published, so you can hear what it actually sounds like before you build anything around it.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production