Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 81 |
dsnote
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation. |
1.6K | 73 | 08/27/26 | MPL-2.0 |
| 82 |
elevenlabs-mcp
The official ElevenLabs MCP server |
1.5K | 250 | 08/20/26 | MIT |
| 83 |
diart
A python package to build AI-powered real-time audio applications |
2.0K | 165 | 06/19/26 | MIT |
| 84 |
audiomentations
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab. |
2.3K | 220 | 04/13/26 | MIT |
| 85 |
amical
🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate. |
1.5K | 139 | 08/26/26 | MIT |
| 86 |
voxtype
Voice-to-text with push-to-talk for Wayland compositors |
1.3K | 100 | 08/29/26 | MIT |
| 87 |
Fun-ASR
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes. |
1.5K | 148 | 08/27/26 | Apache-2.0 |
| 88 |
descript-audio-codec
State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio. |
1.8K | 188 | 07/16/26 | MIT |
| 89 |
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs |
1.5K | 124 | 08/24/26 | Apache-2.0 |
| 90 |
speech-swift
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML |
1.2K | 153 | 08/28/26 | Apache-2.0 |
| 91 |
CrisperWhisper
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps. |
1.4K | 86 | 08/23/26 | Other |
| 92 |
ComfyUI-Qwen-TTS
A Simple Implementation of Qwen3-TTS's ComfyUI |
1.9K | 204 | 06/03/26 | Other |
| 93 |
Whisper-WebUI
A Web UI for easy subtitle using whisper model. |
2.9K | 434 | 12/29/25 | Apache-2.0 |
| 94 |
TTS-Audio-Suite
A ComfyUI custom node suite for multi-engine, multi-language text-to-speech and voice conversion with SRT timing and audio tools. |
1.2K | 140 | 08/28/26 | Other |
| 95 |
MOSS-TTSD
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references. |
1.4K | 136 | 07/26/26 | Apache-2.0 |
| 96 |
whisper-standalone-win
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python. |
3.2K | 165 | 11/07/25 | Other |
| 97 |
StyleTTS2
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models |
6.3K | 700 | 08/10/24 | MIT |
| 98 |
Irene-Voice-Assistant
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины. |
1.2K | 152 | 07/26/26 | Other |
| 99 |
TensorFlowASR
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords |
1.0K | 238 | 08/05/26 | Apache-2.0 |
| 100 |
obs-localvocal
OBS plugin for local speech recognition and captioning using AI |
1.6K | 119 | 05/20/26 | GPL-2.0 |