Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 121 |
DeepDanbooru
AI based multi-label girl image classification system, implemented by using TensorFlow. |
2.9K | 266 | 07/04/26 | MIT |
| 122 |
gpupixel
Real-time image filter engine based on GPU |
2.4K | 342 | 08/29/26 | Apache-2.0 |
| 123 |
manga-ocr
Optical character recognition for Japanese text, with the main focus being Japanese manga |
2.8K | 140 | 07/19/26 | Apache-2.0 |
| 124 |
3dgrut
Ray tracing and hybrid rasterization of Gaussian particles |
2.4K | 279 | 08/19/26 | Apache-2.0 |
| 125 |
RobustVideoMatting
Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML! |
9.5K | 1.2K | 04/02/24 | GPL-3.0 |
| 126 |
emgucv
Emgu CV is a cross platform .Net wrapper to the OpenCV image processing library. |
2.3K | 583 | 08/28/26 | Other |
| 127 |
torch-cam
Class activation maps for your PyTorch models (CAM, Grad-CAM, Grad-CAM++, Smooth Grad-CAM++, Score-CAM, SS-CAM, IS-CAM, XGrad-CAM, Layer-CAM, Finer-CAM, LeGrad, RefineCAM) |
2.3K | 227 | 08/26/26 | Apache-2.0 |
| 128 |
datasets
🎁 7,400,000+ Unsplash images made available for research and machine learning |
2.8K | 142 | 06/26/26 | Other |
| 129 |
BasicSR
Open Source Image and Video Restoration Toolbox for Super-resolution, Denoise, Deblurring, etc. Currently, it includes EDSR, RCAN, SRResNet, SRGAN, ESRGAN, EDVR, BasicVSR, SwinIR, ECBSR, etc. Also support StyleGAN2, DFDNet. |
8.4K | 1.4K | 07/21/24 | Apache-2.0 |
| 130 |
Hierarchical-Localization
Visual localization made easy with hloc |
4.2K | 769 | 12/10/25 | Apache-2.0 |
| 131 |
video-search-and-summarization
GPU-accelerated reference architecture for building video AI agents with real-time verified alerts, visual Q&A, and automated summarization. |
1.8K | 381 | 08/29/26 | Other |
| 132 |
claude-real-video
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT. |
2.1K | 184 | 08/28/26 | MIT |
| 133 |
vipe
ViPE: Video Pose Engine for Geometric 3D Perception |
2.1K | 172 | 08/17/26 | Other |
| 134 |
InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding |
2.4K | 159 | 07/02/26 | Apache-2.0 |
| 135 |
ComfyUI-RMBG
A ComfyUI custom node designed for advanced image background removal and object, face, clothes, and fashion segmentation, utilizing multiple models including RMBG-2.0, INSPYRENET, BEN, BEN2, BiRefNet, SDMatte, SAM, SAM2, SAM3 and GroundingDINO. |
2.1K | 133 | 08/21/26 | GPL-3.0 |
| 136 |
HyperLPR
High Performance Chinese License Plate Recognition Framework. |
6.3K | 2.1K | 02/28/25 | Apache-2.0 |
| 137 |
ocrs
Rust library and CLI tool for OCR (extracting text from images) |
1.9K | 90 | 08/29/26 | Apache-2.0 |
| 138 |
JoyAI-VL-Interaction
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System |
1.8K | 181 | 08/29/26 | Apache-2.0 |
| 139 |
tapnet
Tracking Any Point (TAP) |
2.0K | 187 | 07/22/26 | Apache-2.0 |
| 140 |
Open-Source-Face-Recognition-SDK
The world's 1st open source face recognition SDK for Windows and Linux (Face detection, Face landmark extraction, Face feature extraction, Face template mathcing) |
1.9K | 350 | 08/10/26 | Other |