Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

101–120 of 152

Rank Project Stars Forks
101 MMAudio

[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2.3K 267
102 alltalk_tts

AllTalk TTS is a Coqui-based text-to-speech engine with advanced features like model finetuning, low VRAM support, DeepSpeed acceleration, and JSON API integration for third-party apps.

2.4K 285
103 Speech-AI-Forge

🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

1.4K 189
104 conformer

[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)

1.1K 192
105 IMS-Toucan

Controllable and fast Text-to-Speech for over 7000 languages!

2.2K 316
106 vits

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

7.9K 1.4K
107 FireRedASR

Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.

2.0K 165
108 Whisper-Finetune

Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment

1.2K 223
109 OuteTTS

Interface for OuteTTS models.

1.4K 118
110 index-tts-vllm

Added vLLM support to IndexTTS for faster inference.

1.2K 174
111 VibeVoice-ComfyUI

A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.

1.6K 249
112 whisper-ctranslate2

Whisper command line client compatible with original OpenAI client based on CTranslate2.

1.3K 128
113 distil-whisper

Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.

4.1K 357
114 DeepFilterNet

Noise supression using deep filtering

4.6K 508
115 sherpa-ncnn

Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

1.8K 218
116 pykaldi

A Python wrapper for Kaldi

1.0K 247
117 Whisperboard

The open-source iOS app that's making quality voice transcription more accessible on mobile devices.

1.1K 114
118 vosk-android-demo

Offline speech recognition for Android with Vosk library.

1.1K 277
119 Verbi

A modular voice assistant for experimenting with state-of-the-art transcription, response generation, and text-to-speech models via OpenAI, Groq, Deepgram, and local options.

1.1K 217
120 openai-edge-tts

Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs

2.1K 313