#187 · Primary category: Video & Animation

JoyVASA

audio-driven-talking-face generative-ai image-animation lip-sync portrait-anination talking-head

Diffusion-based Portrait and Animal Animation

Project last updated:04/16/26

GitHub Stars

877

Forks

91

Contributors

3

License

MIT

Why we included this project

JoyVASA turns a single still portrait, or even a photo of an animal, into a talking clip that follows whatever audio you give it. Rather than predicting whole frames, it separates the character's static facial shape from the dynamic expression and head motion, then a diffusion transformer generates that motion straight from the audio. Because the two are decoupled, the same pipeline handles human faces and animal subjects without retraining per identity, and it keeps clips longer and more temporally consistent than many talking-head models. The model trains on a mix of Chinese and English speech, so it copes with both languages reasonably well. Expect a research repo rather than finished software: you'll be downloading checkpoints, getting CUDA set up, and doing the tuning yourself before the output is production-ready.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category