#201 · Primary category: MLOps & Evaluation

Multi-Modality-Arena

chat chatbot chatgpt gradio large-language-models llms multi-modality vision-language-model vqa

Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!

Project last updated:04/21/24

GitHub Stars

566

Forks

39

Contributors

10

License

Other

Why we included this project

Comparing vision-language models fairly is harder than it sounds, and this project tackles that directly. It runs anonymous side-by-side battles on visual question-answering tasks, letting you put models like LLaVA, MiniGPT-4, and BLIP-2 up against each other with images as inputs and collect human preference votes, the same idea behind Chatbot Arena. The repo also ships evaluation code and datasets: the LVLM-eHub benchmark spans dozens of multimodal capability datasets, there's a medical VQA benchmark, and a leaderboard with published scores. So it works both as a quick sanity check on a new model's relative quality and as a reference for building your own evaluation pipeline.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category