Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 121 |
julius
Open-Source Large Vocabulary Continuous Speech Recognition Engine |
1.9K | 304 | 06/16/25 | BSD-3-Clause |
| 122 |
metavoice-src
Foundational model for human-like, expressive TTS |
4.2K | 691 | 07/30/24 | Apache-2.0 |
| 123 |
NeuralNote
Audio Plugin for Audio to MIDI transcription using deep learning. |
2.9K | 190 | 01/16/25 | Apache-2.0 |
| 124 |
vosk-server
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries |
1.3K | 318 | 07/25/25 | Apache-2.0 |
| 125 |
audiolm-pytorch
Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch |
2.6K | 277 | 01/12/25 | MIT |
| 126 |
whisper-jax
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU. |
4.7K | 411 | 04/03/24 | Apache-2.0 |
| 127 |
whisper-web
ML-powered speech recognition directly in your browser |
3.3K | 426 | 10/01/24 | MIT |
| 128 |
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis. |
1.3K | 105 | 06/29/25 | MIT |
| 129 |
soundstorm-pytorch
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch |
1.5K | 94 | 04/24/25 | MIT |
| 130 |
audio-webui
A webui for different audio related Neural Networks |
1.2K | 113 | 05/19/25 | MIT |
| 131 |
autosub
[NO LONGER MAINTAINED] Command-line utility for auto-generating subtitles for any video file |
4.2K | 1.6K | 03/22/24 | MIT |
| 132 |
STT
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy. |
2.6K | 296 | 03/11/24 | MPL-2.0 |
| 133 |
hifi-gan
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis |
2.4K | 557 | 07/27/24 | MIT |
| 134 |
auto-subtitle
Automatically generate and overlay subtitles for any video. |
2.3K | 365 | 07/12/24 | MIT |
| 135 |
masr
中文语音识别; Mandarin Automatic Speech Recognition; |
2.0K | 477 | 07/25/24 | Other |
| 136 |
vocal-remover
Vocal Remover using Deep Neural Networks |
1.8K | 256 | 07/23/24 | MIT |
| 137 |
whisper-writer
💬📝 A small dictation app using OpenAI's Whisper speech recognition model. |
1.1K | 196 | 08/24/24 | GPL-3.0 |
| 138 |
kaldi-gstreamer-server
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork. |
1.1K | 338 | 06/08/24 | BSD-2-Clause |
| 139 |
SpeechT5
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing |
1.4K | 135 | 04/24/24 | MIT |
| 140 |
tensorflow-speech-recognition
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks |
2.2K | 628 | 01/17/24 | Other |