Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 181 |
AdvancedLiterateMachinery
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group. |
1.8K | 197 | 03/17/26 | Apache-2.0 |
| 182 |
fashion-mnist
A MNIST-like fashion product database. Benchmark :point_down: |
12.8K | 3.1K | 06/13/22 | MIT |
| 183 |
nanodet
NanoDet-Plus⚡Super fast and lightweight anchor-free object detection model. 🔥Only 980 KB(int8) / 1.8MB (fp16) and run 97FPS on cellphone🔥 |
6.3K | 1.1K | 08/08/24 | Apache-2.0 |
| 184 |
PointLLM
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds |
1.1K | 58 | 05/15/26 | Other |
| 185 |
ViTPose
The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation" and [TPAMI'23] "ViTPose++: Vision Transformer for Generic Body Pose Estimation" |
2.1K | 268 | 12/25/25 | Apache-2.0 |
| 186 |
unblink
Camera monitoring with VLM |
1.5K | 170 | 03/09/26 | AGPL-3.0 |
| 187 |
OC_SORT
[CVPR2023] The official repo for OC-SORT: Observation-Centric SORT on video Multi-Object Tracking. OC-SORT is simple, online and robust to occlusion/non-linear motion. |
1.1K | 154 | 04/21/26 | MIT |
| 188 |
OmDet
Real-time and accurate open-vocabulary end-to-end object detection |
1.4K | 118 | 03/12/26 | Apache-2.0 |
| 189 |
mmocr
OpenMMLab Text Detection, Recognition and Understanding Toolbox |
4.8K | 781 | 11/27/24 | Apache-2.0 |
| 190 |
ScanNet
ScanNet is a large-scale RGB-D video dataset with 2.5 million views in 1500+ scans, offering 3D poses, reconstructions, and instance-level semantic annotations. |
2.3K | 374 | 11/03/25 | Other |
| 191 |
EasyPR
(CGCSTCD'2017) An easy, flexible, and accurate plate recognition project for Chinese licenses in unconstrained situations. CGCSTCD = China Graduate Contest on Smart-city Technology and Creative Design |
6.4K | 2.5K | 06/04/24 | Apache-2.0 |
| 192 |
layout-parser
A Unified Toolkit for Deep Learning Based Document Image Analysis |
5.8K | 532 | 08/15/24 | Apache-2.0 |
| 193 |
DDColor
[ICCV 2023] DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders |
1.5K | 169 | 01/17/26 | Apache-2.0 |
| 194 |
jepa
PyTorch code and models for V-JEPA self-supervised learning from video. |
4.1K | 419 | 02/27/25 | Other |
| 195 |
CameraView
📸 A well documented, high-level Android interface that makes capturing pictures and videos easy, addressing all of the common issues and needs. Real-time filters, gestures, watermarks, frame processing, RAW, output of any size. |
5.1K | 982 | 09/03/24 | Other |
| 196 |
Video-Depth-Anything
[CVPR 2025 Highlight] Video Depth Anything: Consistent Depth Estimation for Super-Long Videos |
2.1K | 195 | 10/07/25 | Apache-2.0 |
| 197 |
AnimeGANv2
[Open Source]. The improved version of AnimeGAN. Landscape photos/videos to anime |
5.4K | 746 | 08/27/24 | Other |
| 198 |
cleanvision
Automatically find issues in image datasets and practice data-centric computer vision. |
1.2K | 83 | 01/08/26 | Apache-2.0 |
| 199 |
face.evoLVe
🔥🔥High-Performance Face Recognition Library on PaddlePaddle & PyTorch🔥🔥 |
3.6K | 760 | 03/20/25 | MIT |
| 200 |
mmpretrain
OpenMMLab Pre-training Toolbox and Benchmark |
3.9K | 1.1K | 11/01/24 | Apache-2.0 |