Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 61 |
YOLOX
YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with MegEngine, ONNX, TensorRT, ncnn, and OpenVINO supported. Documentation: https://yolox.readthedocs.io/ |
10.6K | 2.5K | 06/08/25 | Apache-2.0 |
| 62 |
Final2x
a cross-platform image super-resolution tool |
7.3K | 525 | 08/20/26 | BSD-3-Clause |
| 63 |
doctr
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k. |
6.3K | 673 | 08/28/26 | Apache-2.0 |
| 64 |
caire
Content aware image resize library |
10.5K | 387 | 05/02/25 | MIT |
| 65 |
opencvsharp
OpenCV wrapper for .NET |
6.1K | 1.2K | 08/24/26 | Apache-2.0 |
| 66 |
jetson-inference
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT and NVIDIA Jetson. |
9.0K | 3.1K | 10/16/25 | MIT |
| 67 |
face-alignment
:fire: 2D and 3D Face alignment library build using pytorch |
7.5K | 1.4K | 04/06/26 | BSD-3-Clause |
| 68 |
DeepLabCut
Official implementation of DeepLabCut: Markerless pose estimation of user-defined features with deep learning for all animals incl. humans |
5.7K | 1.8K | 08/28/26 | LGPL-3.0 |
| 69 |
scrypted
Scrypted is a high performance video integration and automation platform |
5.9K | 391 | 08/29/26 | Other |
| 70 |
PaddleX
All-in-One Development Tool based on PaddlePaddle |
6.3K | 1.2K | 06/25/26 | Apache-2.0 |
| 71 |
SlowFast
PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models. |
7.4K | 1.3K | 03/16/26 | Apache-2.0 |
| 72 |
sahi
Framework agnostic sliced/tiled inference + interactive ui + error analysis plots |
5.5K | 781 | 08/22/26 | MIT |
| 73 |
Pytorch-UNet
PyTorch implementation of the U-Net for image semantic segmentation with high quality images |
11.6K | 2.8K | 08/11/24 | GPL-3.0 |
| 74 |
clip-as-service
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP |
12.8K | 2.1K | 01/23/24 | Other |
| 75 |
Edit-Banana
Edit Banana: A framework for converting statistical formats into editable. |
5.5K | 358 | 08/25/26 | AGPL-3.0 |
| 76 |
facenet
Face recognition using Tensorflow |
14.3K | 4.8K | 07/24/23 | MIT |
| 77 |
remove-ai-watermarks
Remove visible and invisible AI watermarks and provenance metadata from images and video. Python library and CLI for SynthID, C2PA, EXIF, IPTC, XMP, and common generative-AI marks. |
5.3K | 499 | 08/30/26 | Apache-2.0 |
| 78 |
gemini-watermark-remover
A high-performance, 100% client-side tool for removing Gemini AI image & video watermarks. Built with pure JavaScript using mathematically precise Reverse Alpha Blending. / 基于 JavaScript 的纯浏览器端 Gemini AI 图像和视频无损去水印工具,使用数学精确的反向 Alpha 混合算法 |
5.4K | 861 | 08/15/26 | MIT |
| 79 |
sports
computer vision and sports |
5.3K | 656 | 08/28/26 | MIT |
| 80 |
MaaFramework
An automation black-box testing framework based on image recognition |
4.7K | 543 | 08/28/26 | LGPL-3.0 |