#146 · Primary category: MLOps & Evaluation

speech-to-text-benchmark

aws-transcribe cheetah deep-learning deep-neural-networks deepspeech edge-ai google-speech-to-text mozilla-deepspeech offline picovoice pocketsphinx privacy speech-recognition speech-to-text voice-recognition

speech to text benchmark framework

Project last updated:07/08/26

GitHub Stars

697

Forks

72

Contributors

5

License

Apache-2.0

Why we included this project

Choosing a speech-to-text engine usually means trusting vendor benchmarks or running your own ad hoc tests. This framework standardizes the comparison by running the same datasets and metrics across Amazon Transcribe, Azure, Google, IBM Watson, OpenAI Whisper, Whisper.cpp, Vosk, and Picovoice's own engines. It reports word error rate, punctuation error rate, core-hour cost, streaming latency, and model size, and it covers public corpora like LibriSpeech, Common Voice, and TED-LIUM. That mix matters for edge deployments, where a cloud API that wins on accuracy can lose badly on cost or latency. Teams can reuse the published results as a baseline and extend the framework to their own audio and languages.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category