#187 · Primary category: Video & Animation
JoyVASA
Diffusion-based Portrait and Animal Animation
Project last updated:04/16/26
GitHub Stars
877
Forks
91
Contributors
3
License
MIT
Why we included this project
JoyVASA turns a single still portrait, or even a photo of an animal, into a talking clip that follows whatever audio you give it. Rather than predicting whole frames, it separates the character's static facial shape from the dynamic expression and head motion, then a diffusion transformer generates that motion straight from the audio. Because the two are decoupled, the same pipeline handles human faces and animal subjects without retraining per identity, and it keeps clips longer and more temporally consistent than many talking-head models. The model trains on a mix of Chinese and English speech, so it copes with both languages reasonably well. Expect a research repo rather than finished software: you'll be downloading checkpoints, getting CUDA set up, and doing the tuning yourself before the output is production-ready.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
yt-dlp
A feature-rich command-line audio/video downloader
MoneyPrinterTurbo
Generate HD short videos from a topic or keyword with an automated AI workflow.
Deep-Live-Cam
real time face swap and one-click video deepfake with only a single image
manim
Animation engine for explanatory math videos
anime
JavaScript animation engine