#131 · Primary category: Computer Vision
video-search-and-summarization
GPU-accelerated reference architecture for building video AI agents with real-time verified alerts, visual Q&A, and automated summarization.
Project last updated:08/29/26
GitHub Stars
1.8K
Forks
381
Contributors
73
License
Other
Why we included this project
Hours of video become searchable and quotable with this reference architecture, which turns recorded or live footage into something you can interrogate in plain language. NVIDIA packages it as a GPU-accelerated blueprint that wires vision-language models, retrieval over video embeddings, and LLM reasoning into an agent for visual Q&A, long-video summarization, clip retrieval, and verified real-time alerts. It is a template more than a single app, showing how to assemble accelerated microservices from frame and caption ingestion to message-broker publishing, with workflows exposed through the Model Context Protocol. Teams already on NVIDIA hardware or NIM microservices get a runnable starting point instead of gluing together the pieces themselves. If you are building agentic video understanding, this is a concrete implementation worth adapting.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
opencv
Open Source Computer Vision Library
RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
tesseract
Tesseract Open Source OCR Engine (main repository)