#558 · Primary category: Education & Research

dla

deep-learning keyword-spotting signal-processing speaker-verification speech-recognition tts voice-conversion

Deep learning for audio processing

Project last updated:12/15/25

GitHub Stars

762

Forks

123

Contributors

14

License

MIT

Why we included this project

This is the full teaching repository for a university course on deep learning for audio, run at HSE's CS faculty. Instead of a single tool or model, it walks through a structured weekly progression from digital signal processing fundamentals through automatic speech recognition, source separation, audio-visual learning, and text-to-speech, with lectures, seminar notebooks, and homework for each stage. Developers and researchers working on speech and audio ML will find ready-made PyTorch pipelines, experiment-tracking setups, and practical code for CTC and RNN-T training, beam search, and separation architectures like Demucs and ConvTasNet. The course is maintained and re-run each year, so the materials track current practice; that makes them a good self-study path or a base for building your own team onboarding curriculum.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category