#490 · Primary category: AI Agents & Automation

Qwen-MM-Plugins

Make any agent harness multimodal-native.

Project last updated:08/25/26

GitHub Stars

2.8K

Forks

168

Contributors

11

License

Apache-2.0

Why we included this project

Most coding agents are effectively blind: they process text fine, but hand them a screenshot, a video, a PDF, or a 3D model and they have no idea what they are looking at. Qwen-MM-Plugins fixes that by installing multimodal capabilities into the agent harnesses you already use, including Claude Code, Codex, Qoder, OpenClaw, Qwen Code, and Gemini CLI. Each capability is a native Skill with an optional MCP server, so you can add just the pieces you need, whether that is reading images and video, OCR and grounding, web search, or driving Blender and FreeCAD. The guided installer handles the popular CLIs, while manual guides cover less common harnesses, and the core tools run locally without any API key, which makes it easy to test before you commit to cloud credentials.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category