#175 · Primary category: Video & Animation

Sonic

Official implementation of "Sonic: Shifting Focus to Global Audio Perception in Portrait Animation"

Project last updated:01/08/26

GitHub Stars

3.3K

Forks

291

Contributors

5

License

Other

Why we included this project

Sonic starts from a single portrait still and an audio track and produces a lip-synced animation where the head movement comes from the audio itself rather than from reference motion frames. The model reads the global audio context across the whole clip, then splits it into short-range and long-range perception so facial expressions and head pose can be driven independently. In practice that means clips hold up when stretched to several minutes, and the head movements look more natural and varied than in many earlier talking-head models. The CVPR 2025 paper and the reference implementation are a useful anchor for researchers building audio-driven video pipelines. One thing to flag: the license is non-commercial Creative Commons, so this suits academic experimentation far better than commercial product work.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category