#203 · Primary category: Video & Animation

EMO

Emote Portrait Alive: Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Project last updated:08/21/24

GitHub Stars

7.6K

Forks

928

Contributors

3

License

Other

Why we included this project

EMO is a research implementation from Alibaba's Institute for Intelligent Computing, published at ECCV 2024. Give it a single portrait image and an audio track, and it renders a video where the face moves and blinks in time with the speech or singing, with expression changes driven by an audio2video diffusion model that needs only weak conditioning rather than per-subject fine-tuning. The repository carries the official inference code and weights referenced by the paper, so it is a practical starting point for reproducing the results or adapting the approach to your own talking-head or avatar pipeline. The code is research-grade rather than a packaged product, which means you will handle environment setup and write the glue around inference yourself. For researchers and engineers who want a concrete, citable audio2video model to build on, that is a fair trade.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category