Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

21–40 of 152

Rank Project Stars Forks
21 voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

12.7K 1.8K
22 VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

12.0K 1.9K
23 VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

12.0K 1.9K
24 WhisperLiveKit

Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

11.0K 1.1K
25 FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp

11.1K 768
26 moonshine

Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces

11.0K 600
27 so-vits-svc-fork

so-vits-svc fork with realtime support, improved interface and more features.

9.3K 1.2K
28 SenseVoice

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

9.2K 814
29 speech_recognition

Speech recognition module for Python, supporting several engines and APIs, online and offline.

9.0K 2.4K
30 Bert-VITS2

vits2 backbone with multilingual-bert

8.8K 1.3K
31 mlx-audio

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

7.8K 700
32 higgs-audio

Text-audio foundation model from Boson AI

8.3K 640
33 vibe

Transcribe on your own!

7.2K 497
34 ASRT_SpeechRecognition

A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统

8.4K 1.9K
35 annyang

💬 Speech recognition for your site

6.8K 1.0K
36 wav2letter

Facebook AI Research's Automatic Speech Recognition Toolkit

6.4K 988
37 CapsWriter-Offline

PC voice input tool with offline recognition, high accuracy, low latency, hotwords, and LLM polishing; hold CapsLock or mouse side button X2 to speak, release to auto-insert.

6.7K 613
38 pedalboard

🎛 🔊 A Python library for audio.

6.3K 350
39 argmax-oss-swift

On-device Speech AI for Apple Silicon

6.3K 606
40 openwhispr

Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.

5.8K 819