Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
whisper.cpp
Port of OpenAI's Whisper model in C/C++ |
53.3K | 6.1K | 08/29/26 | MIT |
| 2 |
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time |
60.1K | 9.4K | 03/09/26 | Other |
| 3 |
VibeVoice
Open-Source Frontier Voice AI |
53.3K | 6.0K | 07/24/26 | MIT |
| 4 |
voicebox
The open-source AI voice studio. Clone, dictate, create. |
51.8K | 6.5K | 08/09/26 | MIT |
| 5 |
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production |
46.0K | 6.1K | 08/16/24 | MPL-2.0 |
| 6 |
ChatTTS
A generative speech model for daily dialogue. |
39.8K | 4.3K | 04/10/26 | AGPL-3.0 |
| 7 |
MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time |
36.9K | 5.2K | 03/03/26 | Other |
| 8 |
meetily
Privacy-first AI meeting assistant with 4x faster live transcription, speaker diarization, and Ollama summarization, fully local and open-source. |
30.1K | 3.2K | 08/29/26 | MIT |
| 9 |
spleeter
Deezer source separation library including pretrained models. |
28.4K | 3.1K | 06/18/26 | MIT |
| 10 |
whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization) |
23.8K | 2.4K | 07/13/26 | BSD-2-Clause |
| 11 |
faster-whisper
Faster Whisper transcription with CTranslate2 |
25.1K | 2.0K | 11/19/25 | MIT |
| 12 |
CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability. |
23.1K | 2.6K | 05/25/26 | Apache-2.0 |
| 13 |
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech) |
18.4K | 3.6K | 08/29/26 | Apache-2.0 |
| 14 |
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving. |
20.1K | 2.0K | 08/29/26 | MIT |
| 15 |
dia
A TTS model capable of generating ultra-realistic dialogue in one pass. |
19.4K | 1.7K | 11/19/25 | Apache-2.0 |
| 16 |
kaldi
kaldi-asr/kaldi is the official location of the Kaldi project. |
15.5K | 5.4K | 09/22/25 | Other |
| 17 |
vosk-api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node |
15.1K | 1.8K | 08/09/26 | Apache-2.0 |
| 18 |
espnet
End-to-End Speech Processing Toolkit |
9.9K | 2.4K | 08/28/26 | Apache-2.0 |
| 19 |
speechbrain
A PyTorch-based Speech Toolkit |
11.8K | 1.7K | 08/27/26 | Apache-2.0 |
| 20 |
PaddleSpeech
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award. |
12.7K | 2.0K | 08/12/26 | Apache-2.0 |