Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 101 |
MMAudio
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis |
2.3K | 267 | 02/23/26 | MIT |
| 102 |
alltalk_tts
AllTalk TTS is a Coqui-based text-to-speech engine with advanced features like model finetuning, low VRAM support, DeepSpeed acceleration, and JSON API integration for third-party apps. |
2.4K | 285 | 01/09/26 | AGPL-3.0 |
| 103 |
Speech-AI-Forge
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI. |
1.4K | 189 | 05/21/26 | AGPL-3.0 |
| 104 |
conformer
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020) |
1.1K | 192 | 06/29/26 | Apache-2.0 |
| 105 |
IMS-Toucan
Controllable and fast Text-to-Speech for over 7000 languages! |
2.2K | 316 | 01/25/26 | Apache-2.0 |
| 106 |
vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech |
7.9K | 1.4K | 12/06/23 | MIT |
| 107 |
FireRedASR
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability. |
2.0K | 165 | 02/25/26 | Apache-2.0 |
| 108 |
Whisper-Finetune
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment |
1.2K | 223 | 05/08/26 | Apache-2.0 |
| 109 |
OuteTTS
Interface for OuteTTS models. |
1.4K | 118 | 03/23/26 | Apache-2.0 |
| 110 |
index-tts-vllm
Added vLLM support to IndexTTS for faster inference. |
1.2K | 174 | 04/13/26 | Apache-2.0 |
| 111 |
VibeVoice-ComfyUI
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows. |
1.6K | 249 | 02/18/26 | MIT |
| 112 |
whisper-ctranslate2
Whisper command line client compatible with original OpenAI client based on CTranslate2. |
1.3K | 128 | 02/14/26 | MIT |
| 113 |
distil-whisper
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate. |
4.1K | 357 | 01/08/25 | MIT |
| 114 |
DeepFilterNet
Noise supression using deep filtering |
4.6K | 508 | 10/17/24 | Other |
| 115 |
sherpa-ncnn
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc. |
1.8K | 218 | 10/20/25 | Apache-2.0 |
| 116 |
pykaldi
A Python wrapper for Kaldi |
1.0K | 247 | 11/30/25 | Apache-2.0 |
| 117 |
Whisperboard
The open-source iOS app that's making quality voice transcription more accessible on mobile devices. |
1.1K | 114 | 12/18/25 | GPL-3.0 |
| 118 |
vosk-android-demo
Offline speech recognition for Android with Vosk library. |
1.1K | 277 | 12/08/25 | Apache-2.0 |
| 119 |
Verbi
A modular voice assistant for experimenting with state-of-the-art transcription, response generation, and text-to-speech models via OpenAI, Groq, Deepgram, and local options. |
1.1K | 217 | 11/22/25 | MIT |
| 120 |
openai-edge-tts
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs |
2.1K | 313 | 07/01/25 | GPL-3.0 |