#661 · Primary category: AI Agents & Automation

agent-vision-toolkit

agent agent-skills claude-code codex computer-use deepseek dsh-plugin glm harness-engineering multimodal opencode text-only-llm vision vision-language-model

A vision toolkit and skill for text-only LLMs, enabling image Q&A, OCR, UI restoration, and GUI automation with seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode.

Project last updated:08/27/26

GitHub Stars

1.1K

Forks

38

Contributors

7

License

MIT

Why we included this project

Teams that run coding agents on text-only models like DeepSeek often hit a wall the moment an image lands in the conversation, because the model has no eyes of its own. This project moves that capability out of the model and into the harness, shipping command-line tools plus an agent skill that tells the agent which tool to invoke for image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation. A small local proxy and single-file native plugins let pasted images and the built-in image tools work with Codex, Claude Code, Pi, Oh My Pi, and OpenCode without extra prompting. You get to keep the lower cost of a text-only model while the agent still reads full-page screenshots, rebuilds a design into working HTML, or drives a desktop interface step by step. The bundled playbooks capture the actual verification sequence, so the agent follows a proven order of operations instead of improvising from a vague description.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category