#201 · Primary category: MLOps & Evaluation
Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
Project last updated:04/21/24
GitHub Stars
566
Forks
39
Contributors
10
License
Other
Why we included this project
Comparing vision-language models fairly is harder than it sounds, and this project tackles that directly. It runs anonymous side-by-side battles on visual question-answering tasks, letting you put models like LLaVA, MiniGPT-4, and BLIP-2 up against each other with images as inputs and collect human preference votes, the same idea behind Chatbot Arena. The repo also ships evaluation code and datasets: the LVLM-eHub benchmark spans dozens of multimodal capability datasets, there's a medical VQA benchmark, and a leaderboard with published scores. So it works both as a quick sanity check on a new model's relative quality and as a reference for building your own evaluation pipeline.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models