Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 21 |
voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation. |
12.7K | 1.8K | 07/13/26 | GPL-3.0 |
| 22 |
VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. |
12.0K | 1.9K | 08/28/26 | AGPL-3.0 |
| 23 |
VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. |
12.0K | 1.9K | 08/28/26 | AGPL-3.0 |
| 24 |
WhisperLiveKit
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs. |
11.0K | 1.1K | 08/29/26 | Apache-2.0 |
| 25 |
FluidVoice
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp |
11.1K | 768 | 08/29/26 | GPL-3.0 |
| 26 |
moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces |
11.0K | 600 | 08/28/26 | Other |
| 27 |
so-vits-svc-fork
so-vits-svc fork with realtime support, improved interface and more features. |
9.3K | 1.2K | 08/29/26 | Other |
| 28 |
SenseVoice
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection. |
9.2K | 814 | 08/27/26 | MIT |
| 29 |
speech_recognition
Speech recognition module for Python, supporting several engines and APIs, online and offline. |
9.0K | 2.4K | 07/31/26 | BSD-3-Clause |
| 30 |
Bert-VITS2
vits2 backbone with multilingual-bert |
8.8K | 1.3K | 08/24/26 | AGPL-3.0 |
| 31 |
mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon. |
7.8K | 700 | 08/28/26 | MIT |
| 32 |
higgs-audio
Text-audio foundation model from Boson AI |
8.3K | 640 | 06/05/26 | Apache-2.0 |
| 33 |
vibe
Transcribe on your own! |
7.2K | 497 | 08/28/26 | MIT |
| 34 |
ASRT_SpeechRecognition
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统 |
8.4K | 1.9K | 04/10/26 | GPL-3.0 |
| 35 |
annyang
💬 Speech recognition for your site |
6.8K | 1.0K | 08/05/26 | MIT |
| 36 |
wav2letter
Facebook AI Research's Automatic Speech Recognition Toolkit |
6.4K | 988 | 08/28/26 | Other |
| 37 |
CapsWriter-Offline
PC voice input tool with offline recognition, high accuracy, low latency, hotwords, and LLM polishing; hold CapsLock or mouse side button X2 to speak, release to auto-insert. |
6.7K | 613 | 08/26/26 | MIT |
| 38 |
pedalboard
🎛 🔊 A Python library for audio. |
6.3K | 350 | 08/29/26 | GPL-3.0 |
| 39 |
argmax-oss-swift
On-device Speech AI for Apple Silicon |
6.3K | 606 | 08/13/26 | MIT |
| 40 |
openwhispr
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform. |
5.8K | 819 | 08/29/26 | MIT |