Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 201 |
SiamMask
[CVPR19/TPAMI23] SiamMask: A Framework for Fast Online Object Tracking and Segmentation |
3.5K | 804 | 02/14/25 | MIT |
| 202 |
UForm
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️ |
1.2K | 78 | 10/30/25 | Apache-2.0 |
| 203 |
DPED
Software and pre-trained models for automatic photo quality enhancement using Deep Convolutional Networks |
1.7K | 366 | 08/23/25 | Other |
| 204 |
HYPIR
Official implementation of HYPIR: Harnessing Diffusion-Yielded Score Priors for Image Restoration (SIGGRAPH 2025) |
1.2K | 102 | 10/16/25 | Other |
| 205 |
deep-high-resolution-net.pytorch
The project is an official implementation of our CVPR2019 paper "Deep High-Resolution Representation Learning for Human Pose Estimation" |
4.5K | 922 | 08/30/24 | MIT |
| 206 |
BEVFormer
[ECCV 2022] This is the official implementation of BEVFormer, a camera-only framework for autonomous driving perception, e.g., 3D object detection and semantic map segmentation. |
4.6K | 748 | 08/15/24 | Apache-2.0 |
| 207 |
MedSegDiff
Using Diffusion Models to Segment/Reconstruct Organs from Medical Images [AAAI Most influential Paper] |
1.4K | 201 | 09/10/25 | MIT |
| 208 |
pytorch-cifar100
Practice on cifar100(ResNet, DenseNet, VGG, GoogleNet, InceptionV3, InceptionV4, Inception-ResNetv2, Xception, Resnet In Resnet, ResNext,ShuffleNet, ShuffleNetv2, MobileNet, MobileNetv2, SqueezeNet, NasNet, Residual Attention Network, SENet, WideResNet) |
4.8K | 1.2K | 07/15/24 | Other |
| 209 |
ComfyUI-segment-anything-2
ComfyUI nodes to use segment-anything-2 |
1.2K | 87 | 09/28/25 | Apache-2.0 |
| 210 |
Ollama-OCR
OCR package using Ollama vision language models to extract text from images and PDFs, with support for multiple models, output formats, and batch processing. |
2.8K | 323 | 03/17/25 | MIT |
| 211 |
4DGaussians
[CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering |
3.9K | 385 | 10/27/24 | Apache-2.0 |
| 212 |
rPPG-Toolbox
rPPG-Toolbox: Deep Remote PPG Toolbox (NeurIPS 2023) |
1.2K | 304 | 09/15/25 | Other |
| 213 |
monodepth2
[ICCV 2019] Monocular depth estimation from a single image |
4.5K | 987 | 08/10/24 | Other |
| 214 |
Video-LLaVA
【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection |
3.5K | 255 | 12/03/24 | Apache-2.0 |
| 215 |
MAX-Image-Resolution-Enhancer
Upscale an image by a factor of 4, while generating photo-realistic details. |
1.0K | 161 | 09/17/25 | Apache-2.0 |
| 216 |
SimpleCV
The Open Source Framework for Machine Vision |
2.7K | 787 | 12/20/24 | BSD-3-Clause |
| 217 |
SuperGluePretrainedNetwork
SuperGlue: Learning Feature Matching with Graph Neural Networks (CVPR 2020, Oral) |
4.1K | 758 | 08/30/24 | Other |
| 218 |
EasyCV
An all-in-one toolkit for computer vision |
2.0K | 227 | 05/09/25 | Apache-2.0 |
| 219 |
LISA
Project Page for "LISA: Reasoning Segmentation via Large Language Model" |
2.7K | 206 | 02/16/25 | Apache-2.0 |
| 220 |
SimpleCVReproduction
Replication of simple CV Projects including attention, classification, detection, keypoint detection, etc. |
1.3K | 318 | 08/14/25 | Apache-2.0 |