#282 · Primary category: Computer Vision
ComfyUI_Qwen3-VL-Instruct
The successful integration of Qwen3-VL-Instruct series into the ComfyUI platform has enabled a smooth operation, supporting (but not limited to) text-based queries, video queries, single-image queries, and multi-image queries for generating captions or responses.
Project last updated:10/23/25
GitHub Stars
581
Forks
63
Contributors
4
License
Apache-2.0
Why we included this project
ComfyUI users who want vision-language capabilities inside their node graph can install this custom node and get the Qwen3-VL-Instruct models working directly in their workflows. It takes text, single images, multiple images, or video as input, and returns captions or responses, so visual understanding can feed into a larger pipeline instead of living in a separate tool. The bundled example workflows cover all four query modes, and models download automatically on first run, so you don't have to fetch them by hand. Teams that need image captioning or video summarization inside ComfyUI will find this a practical addition. One thing to plan for: you may need the companion Display Text node from the author's MiniCPM-V addon to see outputs.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
opencv
Open Source Computer Vision Library
RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
tesseract
Tesseract Open Source OCR Engine (main repository)