#134 · Primary category: Computer Vision

InternVideo

action-recognition benchmark contrastive-learning foundation-models instruction-tuning masked-autoencoder multimodal open-set-recognition self-supervised spatio-temporal-action-localization temporal-action-localization video-clip video-data video-dataset video-question-answering video-retrieval video-understanding vision-transformer zero-shot-classification zero-shot-retrieval

[ECCV2024] Video Foundation Models & Data for Multimodal Understanding

Project last updated:07/02/26

GitHub Stars

2.4K

Forks

159

Contributors

29

License

Apache-2.0

Why we included this project

InternVideo collects several generations of video foundation models and related work from the OpenGVLab team, all aimed at getting machines to make sense of video rather than just retrieve or replay it. The repo spans action recognition, temporal action localization, video-text retrieval, and open-ended video question answering, so you can find pretrained encoders and chat-style models for a range of understanding tasks. Researchers and ML engineers can pull these checkpoints to build custom video features without training from scratch, and the bundled InternVid video-text dataset is useful for fine-tuning or evaluation. The project is research-oriented, so expect technical reports and model weights more than deployment tooling. Still, the breadth of tasks covered in one place makes it a convenient reference when comparing video backbones.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category