Speech & Audio

ASR, TTS, and speech tooling — alternatives to Whisper API and ElevenLabs.

152 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

61–80 of 152

Rank Project Stars Forks
61 YuE

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

6.4K 757
62 AI-Video-Transcriber

Transcribe and summarize videos and podcasts using AI. Open-source, multi-platform, and supports multiple languages.

3.2K 416
63 TTS-WebUI

A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!

3.2K 329
64 vexa

Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-host or use hosted SaaS.

2.7K 448
65 willow

Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative

3.1K 129
66 whishper

Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!

3.1K 179
67 whisper-timestamped

Multilingual Automatic Speech Recognition with word-level timestamps and confidence

2.8K 211
68 aeneas

aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)

2.9K 275
69 audio.cpp

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.

2.1K 254
70 VieNeu-TTS

Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt

2.4K 732
71 Scriberr

Self-hosted AI audio transcription

3.0K 250
72 wukong-robot

🤖 wukong-robot is a simple, flexible, and elegant Chinese voice dialogue robot/smart speaker project, supporting ChatGPT multi-turn conversation, and possibly the first open-source smart speaker project to support brain-computer interaction.

7.1K 1.4K
73 AudioNotes

Quickly extract audio and video content into a structured markdown note.

2.5K 365
74 asteroid

The PyTorch-based audio source separation toolkit for researchers

2.6K 450
75 WhisperJAV

ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV

2.2K 183
76 audioFlux

A library for audio and music analysis, feature extraction.

3.4K 146
77 lhotse

Tools for handling multimodal data in machine learning projects.

1.1K 277
78 Android-MVVM-Architecture-Android-Voice-AI-SDK

Voice AI SDK is a reusable Android library that gives any app a full voice-driven AI conversation pipeline in minutes. Voice Assistant + Android Voide AI + SDK + MVVM + Kotlin

2.6K 614
79 ElatoAI

Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Companions, and Devices

1.9K 240
80 ClearerVoice-Studio

An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.

4.5K 365