Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

81–100 of 152

Rank Project Stars Forks
81 dsnote

Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.

1.6K 73
82 elevenlabs-mcp

The official ElevenLabs MCP server

1.5K 250
83 diart

A python package to build AI-powered real-time audio applications

2.0K 165
84 audiomentations

A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.

2.3K 220
85 amical

🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.

1.5K 139
86 voxtype

Voice-to-text with push-to-talk for Wayland compositors

1.3K 100
87 Fun-ASR

Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

1.5K 148
88 descript-audio-codec

State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.

1.8K 188
89 SALMONN

SALMONN family: A suite of advanced multi-modal LLMs

1.5K 124
90 speech-swift

AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

1.2K 153
91 CrisperWhisper

Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

1.4K 86
92 ComfyUI-Qwen-TTS

A Simple Implementation of Qwen3-TTS's ComfyUI

1.9K 204
93 Whisper-WebUI

A Web UI for easy subtitle using whisper model.

2.9K 434
94 TTS-Audio-Suite

A ComfyUI custom node suite for multi-engine, multi-language text-to-speech and voice conversion with SRT timing and audio tools.

1.2K 140
95 MOSS-TTSD

MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.

1.4K 136
96 whisper-standalone-win

Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.

3.2K 165
97 StyleTTS2

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

6.3K 700
98 Irene-Voice-Assistant

Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.

1.2K 152
99 TensorFlowASR

:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords

1.0K 238
100 obs-localvocal

OBS plugin for local speech recognition and captioning using AI

1.6K 119