Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 241 |
pytorch-segmentation
:art: Semantic segmentation models, datasets and losses implemented in PyTorch. |
1.8K | 389 | 02/23/25 | MIT |
| 242 |
tensorflow-yolov3
🔥 TensorFlow Code for technical report: "YOLOv3: An Incremental Improvement" |
3.6K | 1.3K | 07/04/24 | MIT |
| 243 |
SegmentAnythingin3D
Segment Anything in 3D with NeRFs (NeurIPS 2023 & IJCV 2025) |
1.0K | 64 | 05/19/25 | Apache-2.0 |
| 244 |
ghost
A new one shot face swap approach for image and video domains |
1.6K | 308 | 02/25/25 | Apache-2.0 |
| 245 |
DenseNet
Densely Connected Convolutional Networks, In CVPR 2017 (Best Paper Award). |
4.9K | 1.1K | 01/09/24 | BSD-3-Clause |
| 246 |
PyTorch-Encoding
A CV toolkit for my papers. |
2.0K | 448 | 12/21/24 | MIT |
| 247 |
mAP
mean Average Precision - This code evaluates the performance of your neural net for object recognition. |
3.0K | 918 | 08/15/24 | Apache-2.0 |
| 248 |
DIS
This is the repo for our new project Highly Accurate Dichotomous Image Segmentation |
2.6K | 289 | 09/23/24 | Apache-2.0 |
| 249 |
DeepMosaics
Automatically remove the mosaics in images and videos, or add mosaics to them. |
2.6K | 486 | 08/30/24 | GPL-3.0 |
| 250 |
DINO
[ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection" |
2.8K | 309 | 07/31/24 | Apache-2.0 |
| 251 |
XMem
[ECCV 2022] XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model |
2.0K | 212 | 11/15/24 | MIT |
| 252 |
yolov3-tf2
YoloV3 Implemented in Tensorflow 2.0 |
2.5K | 888 | 08/30/24 | MIT |
| 253 |
deep-text-recognition-benchmark
Text recognition (optical character recognition) with deep learning methods, ICCV 2019 |
3.9K | 1.1K | 03/04/24 | Apache-2.0 |
| 254 |
ScaledYOLOv4
Scaled-YOLOv4: Scaling Cross Stage Partial Network |
2.0K | 563 | 11/03/24 | GPL-3.0 |
| 255 |
hagrid
HAnd Gesture Recognition Image Dataset |
1.1K | 149 | 02/27/25 | Other |
| 256 |
Video-LLaMA
[EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding |
3.1K | 287 | 06/04/24 | BSD-3-Clause |
| 257 |
DAMO-YOLO
DAMO-YOLO: a fast and accurate object detection method with some new techs, including NAS backbones, efficient RepGFPN, ZeroHead, AlignedOTA, and distillation enhancement. |
3.2K | 407 | 05/25/24 | Apache-2.0 |
| 258 |
Sophus
C++ implementation of Lie Groups using Eigen. |
2.4K | 671 | 07/06/24 | Other |
| 259 |
decord
An efficient video loader for deep learning with smart shuffling that's super easy to digest |
2.5K | 233 | 07/17/24 | Apache-2.0 |
| 260 |
darknet_ros
YOLO ROS: Real-Time Object Detection for ROS |
2.4K | 1.2K | 07/19/24 | BSD-3-Clause |