Computer Vision
Detection, segmentation, OCR, and vision pipelines — production CV open source.
384 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 81 |
sparrow
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM |
5.2K | 519 | 08/26/26 | GPL-3.0 |
| 82 |
openMVG
open Multiple View Geometry library. Basis for 3D computer vision and Structure from Motion. |
6.5K | 1.7K | 12/17/25 | MPL-2.0 |
| 83 |
watermark-removal
a machine learning image inpainting task that instinctively removes watermarks from image indistinguishable from the ground truth image |
5.1K | 599 | 08/14/26 | Other |
| 84 |
trace.moe
Timestamp Retrieval for Anime Clips Everywhere |
5.0K | 263 | 08/21/26 | MIT |
| 85 |
nsfw_data_scraper
Collection of scripts to aggregate image data for the purposes of training an NSFW Image Classifier |
12.6K | 2.8K | 01/21/24 | MIT |
| 86 |
torchgeo
TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data |
4.2K | 581 | 08/29/26 | MIT |
| 87 |
Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation. |
6.0K | 552 | 03/31/26 | MIT |
| 88 |
AliceVision
3D Computer Vision Framework |
3.5K | 878 | 08/28/26 | Other |
| 89 |
obs-backgroundremoval
An OBS plugin for removing background in portrait images (video), making it easy to replace the background when recording or streaming. |
4.5K | 288 | 08/29/26 | GPL-3.0 |
| 90 |
scenic
Scenic: A Jax Library for Computer Vision Research and Beyond |
3.8K | 479 | 08/28/26 | Apache-2.0 |
| 91 |
lightly
A python library for self-supervised learning on images. |
3.8K | 356 | 08/28/26 | MIT |
| 92 |
segment-geospatial
A Python package for segmenting geospatial data with the Segment Anything Model (SAM) |
4.1K | 440 | 08/24/26 | MIT |
| 93 |
super-gradients
Easily train or fine-tune SOTA computer vision models with one open source training library. The home of Yolo-NAS. |
5.1K | 592 | 02/24/26 | Apache-2.0 |
| 94 |
modlens
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). |
3.8K | 110 | 08/27/26 | MIT |
| 95 |
vision-agent
This tool has been deprecated. Use Agentic Document Extraction instead. |
5.3K | 601 | 01/29/26 | Apache-2.0 |
| 96 |
co-tracker
CoTracker is a model for tracking any point (pixel) on a video. |
5.1K | 386 | 03/03/26 | Other |
| 97 |
U-2-Net
The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection." |
9.9K | 1.6K | 06/26/24 | Apache-2.0 |
| 98 |
nunif
Misc; latest version of waifu2x; 2D video to stereo 3D video conversion |
3.4K | 282 | 08/22/26 | MIT |
| 99 |
geoai
GeoAI: Artificial Intelligence for Geospatial Data |
3.3K | 472 | 08/24/26 | MIT |
| 100 |
vjepa2
PyTorch code and models for VJEPA2 self-supervised learning from video. |
4.5K | 559 | 03/23/26 | MIT |