Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 161 |
MaterialSearch
Semantic search. Search local photos and videos through natural language. AI语义搜索本地素材。以图搜图、查找本地素材、根据文字描述匹配画面、视频帧搜索、根据画面描述搜索视频。 |
2.0K | 217 | 05/23/26 | GPL-3.0 |
| 162 |
pysot
SenseTime Research platform for single object tracking, implementing algorithms like SiamRPN and SiamMask. |
4.6K | 1.1K | 06/22/25 | Apache-2.0 |
| 163 |
CityGaussian
[ECCV`24&ICLR`25] CityGaussian Series for High-quality Large-Scale Scene Reconstruction with Gaussians |
1.3K | 110 | 08/16/26 | Other |
| 164 |
claude-video-vision
Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis |
1.3K | 152 | 08/07/26 | MIT |
| 165 |
text-extract-api
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown |
3.2K | 278 | 12/08/25 | MIT |
| 166 |
withoutbg-python
Python SDK for local and cloud background removal (pip install withoutbg) |
1.3K | 62 | 07/23/26 | Other |
| 167 |
Objectron
A dataset of 15K annotated object-centric video clips with 3D bounding boxes, camera poses, and AR metadata for 3D object detection. |
2.3K | 266 | 03/06/26 | Other |
| 168 |
byol-pytorch
Usable Implementation of "Bootstrap Your Own Latent" self-supervised learning, from Deepmind, in Pytorch |
1.9K | 248 | 04/27/26 | MIT |
| 169 |
pigo
Fast face detection, pupil/eyes localization and facial landmark points detection library in pure Go. |
4.7K | 325 | 05/03/25 | MIT |
| 170 |
ultralyticsPro
🔥🔥🔥 专注于YOLO11,YOLOv8、TYOLOv12、YOLOv10、RT-DETR、YOLOv7、YOLOv5改进模型,Support to improve backbone, neck, head, loss, IoU, NMS and other modules🚀 |
3.0K | 458 | 12/15/25 | Other |
| 171 |
glue-factory
Training library for local feature detection and matching |
1.1K | 157 | 07/20/26 | Apache-2.0 |
| 172 |
MambaVision
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone |
2.2K | 150 | 03/11/26 | Other |
| 173 |
opendatacam
An open source tool to quantify the world |
1.7K | 300 | 04/23/26 | MIT |
| 174 |
BackgroundMattingV2
Real-Time High-Resolution Background Matting |
7.2K | 960 | 06/19/24 | MIT |
| 175 |
zed-sdk
⚡️The spatial perception framework for rapidly building smart robots and spaces |
1.2K | 507 | 06/18/26 | MIT |
| 176 |
jeelizFaceFilter
🎭 Lightweight WebGL & JavaScript library for real-time multi-face detection, tracking and augmented-reality face filters. Supports rotation, mouth-opening and integrates with Three.js, Babylon.js, Canvas2D and CSS3D |
2.9K | 546 | 11/14/25 | Apache-2.0 |
| 177 |
deepseek_ocr_app
A quick vibe coded app for deepseek OCR |
1.9K | 310 | 03/31/26 | MIT |
| 178 |
CVprojects
computer vision projects | 计算机视觉相关好玩的AI项目(Python、C++、embedded system) |
2.6K | 692 | 12/26/25 | Other |
| 179 |
donut
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022 |
6.9K | 565 | 07/11/24 | MIT |
| 180 |
VGGT-SLAM
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold |
1.1K | 119 | 06/29/26 | BSD-2-Clause |