#128 · Primary category: Speech & Audio
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
Project last updated:06/29/25
GitHub Stars
1.3K
Forks
105
Contributors
2
License
MIT
Why we included this project
StreamSpeech takes a single trained model and covers offline and streaming speech recognition, speech-to-text translation, speech-to-speech translation, and text-to-speech synthesis, so you don't have to bolt together separate tools for each step. During simultaneous translation it can surface intermediate ASR transcripts and text translations as they are produced, which makes it useful for live captions or draft subtitles before the spoken output finishes. Pretrained French, Spanish, and German-to-English models are on Hugging Face, and the repo bundles the ACL 2024 paper code, training recipes, and comparison tools. Teams prototyping real-time interpretation can also try the web demo or local GUI before integrating the model into their own stack.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production