#91 · Primary category: Speech & Audio

CrisperWhisper

asr audio detection filler full-duplex-audio labeling recognition speech speech-processing speech-recognition stutter-detection stuttering timestamps transcription turn-taking verbatim whisper

Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

Project last updated:08/23/26

GitHub Stars

1.4K

Forks

86

Contributors

5

License

Other

Why we included this project

Most transcription tools quietly smooth over what people actually say, dropping the ums and false starts in favor of a 'clean' transcript. CrisperWhisper works the other way: it can transcribe word for word with fillers and vocal events marked, or produce an 'intended' version where numbers, dates, and emails are formatted the way you'd write them. Word-level timestamps land around 30 to 40 ms of boundary error, tight enough for alignment work. The Verbatimize mode is the standout for teams that already hold trusted clean transcripts: feed it the original audio plus that clean text and it reproduces the spoken disfluencies word for word, which matters when you're preparing TTS or clinical speech-analysis data, or turning existing clean corpora into verbatim datasets. It covers most languages Whisper supports and handles long recordings without the usual chunk-boundary artifacts.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category