#177 · Primary category: Computer Vision

dsh-vision-toolkit

agent-skills agent-vision-toolkit computer-vision deepseek deepseek-harness dsh dsh-plugin gui-automation ocr plugin python screenshot-testing text-only-llm typescript ui-restoration vision-language-model vision-tools

[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

Project last updated:08/29/26

GitHub Stars

842

Forks

38

Contributors

6

License

MIT

Why we included this project

Text-only models go blind the moment a task involves a screenshot, a scanned page, or a UI mockup. This toolkit gives DeepSeek Harness agents a set of callable vision skills instead: paste an image and ask questions about it, run OCR across long screenshots, turn a screenshot into front-end code, and compare UI states with pixel-level diffs. A single command installs it, and it runs both in the web interface and headless, so the same setup serves interactive use and automated pipelines. Teams building GUI automation, visual regression checks, or document-processing agents on text-only models get a practical way to add sight without changing the model underneath.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category