#135 · Primary category: Knowledge Base & RAG
VLM2Vec
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
Project last updated:08/23/26
GitHub Stars
680
Forks
64
Contributors
10
License
Apache-2.0
Why we included this project
Most embedding tooling assumes your data is text. VLM2Vec is the reference implementation for the other case: adapting a vision-language model into a dense retriever that handles images, video, audio, visual documents, and agent-centric queries alongside text. It ships the training recipe plus the MMEB benchmark, now at V3 with 190 tasks, which scores how well any embedding model follows modality-specific retrieval instructions. That makes it useful both for training your own model and for evaluating third-party ones before you commit. The OmniSET diagnostic component helps separate failures caused by modality handling from failures in instruction following. It is research code first, so plan to adapt the scripts rather than drop them into production unchanged.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ragflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Understand-Anything
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
crawl4ai
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
docling
Get your documents ready for gen AI
anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience