#144 · Primary category: MLOps & Evaluation
MMMU
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
Project last updated:07/28/26
GitHub Stars
593
Forks
56
Contributors
8
License
Apache-2.0
Why we included this project
MMMU is a demanding public benchmark for anyone training or fine-tuning multimodal models who wants to know whether the model actually reasons rather than pattern-matches. Its roughly 11,500 college-level questions come from exams, quizzes, and textbooks across six disciplines, and the images span everything from charts and diagrams to music sheets and chemical structures, so it tests perception and domain knowledge together. The repo includes the evaluation code and, since early 2026, the test-set answers, which means you can score your own model locally instead of depending on a hosted leaderboard. The MMMU-Pro companion adds harder multiple-choice items and vision-only settings for a stricter read on genuine understanding. For teams comparing open-weight vision-language models, this is a practical, reproducible yardstick.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models