#139 · Primary category: Video & Animation

echomimic_v2

audio-driven-body-animation audio-driven-portrait-animations audio-driven-talking-face cvpr2025 human-animation talking-face-generation talking-head video-generation

[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

Project last updated:02/23/26

GitHub Stars

4.7K

Forks

551

Contributors

6

License

Apache-2.0

Why we included this project

Point a single photo at an audio track and EchoMimicV2 turns it into a person who speaks and gestures. It is semi-body animation, so lips, expression, head motion, and upper-body gestures all follow the audio rather than just the face moving. The code comes from Ant Group's research team, was accepted at CVPR 2025, and ships pretrained weights for English and Mandarin plus an accelerated single-GPU inference path. Gradio and ComfyUI interfaces let you try it without wiring up a pipeline yourself, though this is research software: expect to install dependencies, fetch checkpoints, and tune configs. That effort pays off for teams prototyping talking-person videos for content or avatar demos, and the EMTD dataset processing scripts give researchers a workable base for building on audio-driven animation.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category