#91 · Primary category: Speech & Audio
CrisperWhisper
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
Project last updated:08/23/26
GitHub Stars
1.4K
Forks
86
Contributors
5
License
Other
Why we included this project
Most transcription tools quietly smooth over what people actually say, dropping the ums and false starts in favor of a 'clean' transcript. CrisperWhisper works the other way: it can transcribe word for word with fillers and vocal events marked, or produce an 'intended' version where numbers, dates, and emails are formatted the way you'd write them. Word-level timestamps land around 30 to 40 ms of boundary error, tight enough for alignment work. The Verbatimize mode is the standout for teams that already hold trusted clean transcripts: feed it the original audio plus that clean text and it reproduces the spoken disfluencies word for word, which matters when you're preparing TTS or clinical speech-analysis data, or turning existing clean corpora into verbatim datasets. It covers most languages Whisper supports and handles long recordings without the usual chunk-boundary artifacts.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
whisper.cpp
Port of OpenAI's Whisper model in C/C++
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VibeVoice
Open-Source Frontier Voice AI
voicebox
The open-source AI voice studio. Clone, dictate, create.
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production