#128 · Primary category: Speech & Audio

StreamSpeech

all-in-one asr audio-processing machine-translation non-autoregressive seamless simultaneous-translation speech speech-enhancement speech-processing speech-recognition speech-synthesis speech-to-text speech-translation streaming-audio text-to-audio text-to-speech translation tts voice

StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

Project last updated:06/29/25

GitHub Stars

1.3K

Forks

105

Contributors

2

License

MIT

Why we included this project

StreamSpeech takes a single trained model and covers offline and streaming speech recognition, speech-to-text translation, speech-to-speech translation, and text-to-speech synthesis, so you don't have to bolt together separate tools for each step. During simultaneous translation it can surface intermediate ASR transcripts and text translations as they are produced, which makes it useful for live captions or draft subtitles before the spoken output finishes. Pretrained French, Spanish, and German-to-English models are on Hugging Face, and the repo bundles the ACL 2024 paper code, training recipes, and comparison tools. Teams prototyping real-time interpretation can also try the web demo or local GUI before integrating the model into their own stack.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category