Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 61 |
YuE
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open |
6.4K | 757 | 06/04/25 | Apache-2.0 |
| 62 |
AI-Video-Transcriber
Transcribe and summarize videos and podcasts using AI. Open-source, multi-platform, and supports multiple languages. |
3.2K | 416 | 08/23/26 | Apache-2.0 |
| 63 |
TTS-WebUI
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark! |
3.2K | 329 | 07/27/26 | MIT |
| 64 |
vexa
Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-host or use hosted SaaS. |
2.7K | 448 | 08/29/26 | Apache-2.0 |
| 65 |
willow
Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative |
3.1K | 129 | 08/04/26 | Apache-2.0 |
| 66 |
whishper
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models! |
3.1K | 179 | 07/31/26 | AGPL-3.0 |
| 67 |
whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence |
2.8K | 211 | 08/17/26 | AGPL-3.0 |
| 68 |
aeneas
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment) |
2.9K | 275 | 07/25/26 | AGPL-3.0 |
| 69 |
audio.cpp
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency. |
2.1K | 254 | 08/29/26 | Other |
| 70 |
VieNeu-TTS
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt |
2.4K | 732 | 08/25/26 | Apache-2.0 |
| 71 |
Scriberr
Self-hosted AI audio transcription |
3.0K | 250 | 06/01/26 | MIT |
| 72 |
wukong-robot
🤖 wukong-robot is a simple, flexible, and elegant Chinese voice dialogue robot/smart speaker project, supporting ChatGPT multi-turn conversation, and possibly the first open-source smart speaker project to support brain-computer interaction. |
7.1K | 1.4K | 10/25/24 | MIT |
| 73 |
AudioNotes
Quickly extract audio and video content into a structured markdown note. |
2.5K | 365 | 08/19/26 | MIT |
| 74 |
asteroid
The PyTorch-based audio source separation toolkit for researchers |
2.6K | 450 | 05/13/26 | MIT |
| 75 |
WhisperJAV
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV |
2.2K | 183 | 08/22/26 | MIT |
| 76 |
audioFlux
A library for audio and music analysis, feature extraction. |
3.4K | 146 | 03/06/26 | MIT |
| 77 |
lhotse
Tools for handling multimodal data in machine learning projects. |
1.1K | 277 | 08/26/26 | Apache-2.0 |
| 78 |
Android-MVVM-Architecture-Android-Voice-AI-SDK
Voice AI SDK is a reusable Android library that gives any app a full voice-driven AI conversation pipeline in minutes. Voice Assistant + Android Voide AI + SDK + MVVM + Kotlin |
2.6K | 614 | 06/01/26 | Other |
| 79 |
ElatoAI
Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Companions, and Devices |
1.9K | 240 | 08/19/26 | Other |
| 80 |
ClearerVoice-Studio
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc. |
4.5K | 365 | 08/14/25 | Apache-2.0 |