Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

1–20 of 152

Rank Project Stars Forks
1 whisper.cpp

Port of OpenAI's Whisper model in C/C++

53.3K 6.1K
2 Real-Time-Voice-Cloning

Clone a voice in 5 seconds to generate arbitrary speech in real-time

60.1K 9.4K
3 VibeVoice

Open-Source Frontier Voice AI

53.3K 6.0K
4 voicebox

The open-source AI voice studio. Clone, dictate, create.

51.8K 6.5K
5 TTS

🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

46.0K 6.1K
6 ChatTTS

A generative speech model for daily dialogue.

39.8K 4.3K
7 MockingBird

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

36.9K 5.2K
8 meetily

Privacy-first AI meeting assistant with 4x faster live transcription, speaker diarization, and Ollama summarization, fully local and open-source.

30.1K 3.2K
9 spleeter

Deezer source separation library including pretrained models.

28.4K 3.1K
10 whisperX

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

23.8K 2.4K
11 faster-whisper

Faster Whisper transcription with CTranslate2

25.1K 2.0K
12 CosyVoice

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

23.1K 2.6K
13 Speech

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

18.4K 3.6K
14 FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

20.1K 2.0K
15 dia

A TTS model capable of generating ultra-realistic dialogue in one pass.

19.4K 1.7K
16 kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

15.5K 5.4K
17 vosk-api

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

15.1K 1.8K
18 espnet

End-to-End Speech Processing Toolkit

9.9K 2.4K
19 speechbrain

A PyTorch-based Speech Toolkit

11.8K 1.7K
20 PaddleSpeech

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

12.7K 2.0K
< Previous
/ 8
Next >