Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

141–152 of 152

Rank Project Stars Forks
141 whisper-turbo

Cross-Platform, GPU Accelerated Whisper 🏎️

1.8K 82
142 deepvoice3_pytorch

PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models

2.0K 481
143 tacotron

A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)

3.0K 936
144 Automatic_Speech_Recognition

End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow

2.8K 535
145 audio-diffusion-pytorch

Audio generation using diffusion models, in PyTorch.

2.1K 175
146 naturalspeech2-pytorch

Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch

1.3K 104
147 Speech-Emotion-Analyzer

The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)

1.4K 435
148 artyom.js

A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within your website.

1.3K 362
149 pytorch-kaldi

pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.

2.4K 442
150 DeepAudioClassification

Finding the genre of a song with Deep Learning

1.1K 217
151 SincNet

SincNet is a neural architecture for efficiently processing raw audio samples.

1.2K 272
152 project_alias

A teachable parasite that lets users customize wake-words and commands for smart assistants, enhancing privacy and control via a Raspberry Pi-based device.

1.7K 98
< Previous
/ 8
Next >