#106 · Primary category: Computer Vision
Ask-Anything
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
Project last updated:07/17/26
GitHub Stars
3.3K
Forks
268
Contributors
13
License
MIT
Why we included this project
Ask-Anything collects several related video-language models built around one idea: you talk to a model about what's happening in a video or an image instead of just getting a caption back. The project grew out of the VideoChat line, which earned a CVPR 2024 highlight for treating video understanding as a chat task, and it now offers both end-to-end trained models and variants that use ChatGPT, StableLM, MOSS, or MiniGPT-4 as the language half. Beyond the pretrained weights, the repo carries training and instruction-tuning pipelines, released instruction datasets, and demo notebooks, so it doubles as a reference for teams building video Q&A features or summarizing footage. That makes it useful for anyone who wants to compare open video-language models or fine-tune one toward a specific task without starting from scratch.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
opencv
Open Source Computer Vision Library
RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
tesseract
Tesseract Open Source OCR Engine (main repository)