#101 · Primary category: Speech & Audio

MMAudio

audio audio-synthesis computer-vision deep-learning text-to-audio video-to-audio

[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Project last updated:02/23/26

GitHub Stars

2.3K

Forks

267

Contributors

3

License

MIT

Why we included this project

MMAudio, from a CVPR 2025 paper, turns a video clip, optionally paired with a text prompt, into audio that stays synchronized with the frames. A dedicated synchronization module keeps the generated sound aligned to the visuals, which makes it useful for adding sound to silent stock footage or to videos generated by tools like Sora and Veo that don't produce audio. The repo ships pretrained weights plus Colab and Replicate demos, so you can judge the output without setting up a local environment. If you regularly produce short-form video and want to avoid closed APIs for sound design, this is worth a look.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category