Speech & Audio
ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.
152 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 41 |
wenet
Production First and Production Ready End-to-End Speech Recognition Toolkit |
5.2K | 1.2K | 06/15/26 | Apache-2.0 |
| 42 |
abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions. |
5.8K | 432 | 08/29/26 | MIT |
| 43 |
whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper |
5.6K | 503 | 08/15/26 | BSD-2-Clause |
| 44 |
porcupine
On-device wake word detection powered by deep learning |
4.9K | 577 | 08/12/26 | Apache-2.0 |
| 45 |
audio
Data manipulation and transformation for audio signal processing, powered by PyTorch |
2.9K | 795 | 08/29/26 | BSD-2-Clause |
| 46 |
Applio
A simple, high-quality voice conversion tool focused on ease of use and performance. |
3.7K | 587 | 08/28/26 | MIT |
| 47 |
SmartSub
Free open-source desktop app for video subtitles: local Whisper speech-to-text, translation, AI dubbing with voice cloning, and subtitle burning, GPU-accelerated across Windows/macOS/Linux. |
4.8K | 351 | 08/20/26 | MIT |
| 48 |
pocketsphinx
A small speech recognizer |
4.3K | 732 | 08/26/26 | Other |
| 49 |
Orpheus-TTS
Towards Human-Sounding Speech |
6.3K | 535 | 12/05/25 | Apache-2.0 |
| 50 |
MOSS-TTS
Open-source family of high-fidelity speech and sound generation models supporting long-form TTS, multi-speaker dialogue, voice design, sound effects, and real-time streaming. |
4.0K | 366 | 07/26/26 | Apache-2.0 |
| 51 |
openless
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. |
3.4K | 298 | 08/27/26 | MIT |
| 52 |
pyAudioAnalysis
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications |
6.3K | 1.2K | 08/04/25 | Apache-2.0 |
| 53 |
basic-pitch
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection |
5.5K | 502 | 11/13/25 | Apache-2.0 |
| 54 |
elevenlabs-python
The official Python SDK for the ElevenLabs API. |
3.1K | 439 | 08/25/26 | MIT |
| 55 |
whisper-asr-webservice
OpenAI Whisper ASR Webservice API |
3.3K | 584 | 08/09/26 | MIT |
| 56 |
fastrtc
The python library for real-time communication |
4.6K | 433 | 01/12/26 | MIT |
| 57 |
stt
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式 |
4.8K | 499 | 01/22/26 | GPL-3.0 |
| 58 |
TTS
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts) |
10.2K | 1.3K | 11/09/23 | MPL-2.0 |
| 59 |
EmotiVoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine |
8.5K | 756 | 08/13/24 | Apache-2.0 |
| 60 |
MeloTTS
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean. |
7.6K | 1.1K | 12/24/24 | MIT |