Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

121–140 of 152

Rank Project Stars Forks
121 julius

Open-Source Large Vocabulary Continuous Speech Recognition Engine

1.9K 304
122 metavoice-src

Foundational model for human-like, expressive TTS

4.2K 691
123 NeuralNote

Audio Plugin for Audio to MIDI transcription using deep learning.

2.9K 190
124 vosk-server

WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries

1.3K 318
125 audiolm-pytorch

Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch

2.6K 277
126 whisper-jax

JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.

4.7K 411
127 whisper-web

ML-powered speech recognition directly in your browser

3.3K 426
128 StreamSpeech

StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

1.3K 105
129 soundstorm-pytorch

Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch

1.5K 94
130 audio-webui

A webui for different audio related Neural Networks

1.2K 113
131 autosub

[NO LONGER MAINTAINED] Command-line utility for auto-generating subtitles for any video file

4.2K 1.6K
132 STT

🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.

2.6K 296
133 hifi-gan

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

2.4K 557
134 auto-subtitle

Automatically generate and overlay subtitles for any video.

2.3K 365
135 masr

中文语音识别; Mandarin Automatic Speech Recognition;

2.0K 477
136 vocal-remover

Vocal Remover using Deep Neural Networks

1.8K 256
137 whisper-writer

💬📝 A small dictation app using OpenAI's Whisper speech recognition model.

1.1K 196
138 kaldi-gstreamer-server

Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.

1.1K 338
139 SpeechT5

Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

1.4K 135
140 tensorflow-speech-recognition

🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks

2.2K 628