#134 · Primary category: Computer Vision
InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
Project last updated:07/02/26
GitHub Stars
2.4K
Forks
159
Contributors
29
License
Apache-2.0
Why we included this project
InternVideo collects several generations of video foundation models and related work from the OpenGVLab team, all aimed at getting machines to make sense of video rather than just retrieve or replay it. The repo spans action recognition, temporal action localization, video-text retrieval, and open-ended video question answering, so you can find pretrained encoders and chat-style models for a range of understanding tasks. Researchers and ML engineers can pull these checkpoints to build custom video features without training from scratch, and the bundled InternVid video-text dataset is useful for fine-tuning or evaluation. The project is research-oriented, so expect technical reports and model weights more than deployment tooling. Still, the breadth of tasks covered in one place makes it a convenient reference when comparing video backbones.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
opencv
Open Source Computer Vision Library
RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
tesseract
Tesseract Open Source OCR Engine (main repository)