#95 · Primary category: Speech & Audio

MOSS-TTSD

finetune large-language-models sglang speech-dialogue-generation streaming text-to-speeh

MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.

Project last updated:07/26/26

GitHub Stars

1.4K

Forks

136

Contributors

10

License

Apache-2.0

Why we included this project

Most TTS models give you one clean voice reading one passage, which is why they fall apart on anything that sounds like two people talking. MOSS-TTSD is built for the other case: it takes a script with multiple speakers and produces a continuous, emotionally consistent conversation, and it keeps that up over long stretches. You can steer who is speaking, it works across several languages, and the zero-shot cloning means a short audio clip is enough to reproduce a known voice for a new character, no per-speaker training required. If you are producing audiobooks, dubbing, sports commentary, or any dialogue-heavy content, this is worth a look. The model weights and a live demo space are both published, so you can hear what it actually sounds like before you build anything around it.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category