#177 · Primary category: Computer Vision
dsh-vision-toolkit
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
Project last updated:08/29/26
GitHub Stars
842
Forks
38
Contributors
6
License
MIT
Why we included this project
Text-only models go blind the moment a task involves a screenshot, a scanned page, or a UI mockup. This toolkit gives DeepSeek Harness agents a set of callable vision skills instead: paste an image and ask questions about it, run OCR across long screenshots, turn a screenshot into front-end code, and compare UI states with pixel-level diffs. A single command installs it, and it runs both in the web interface and headless, so the same setup serves interactive use and automated pipelines. Teams building GUI automation, visual regression checks, or document-processing agents on text-only models get a practical way to add sight without changing the model underneath.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
opencv
Open Source Computer Vision Library
RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
tesseract
Tesseract Open Source OCR Engine (main repository)