#144 · Primary category: MLOps & Evaluation

MMMU

computer-vision deep-learning deep-neural-networks evaluation foundation-models large-language-models large-multimodal-models llm llms machine-learning multimodal multimodal-deep-learning multimodal-learning multimodality natural-language-processing question-answering stem visual-question-answering

This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

Project last updated:07/28/26

GitHub Stars

593

Forks

56

Contributors

8

License

Apache-2.0

Why we included this project

MMMU is a demanding public benchmark for anyone training or fine-tuning multimodal models who wants to know whether the model actually reasons rather than pattern-matches. Its roughly 11,500 college-level questions come from exams, quizzes, and textbooks across six disciplines, and the images span everything from charts and diagrams to music sheets and chemical structures, so it tests perception and domain knowledge together. The repo includes the evaluation code and, since early 2026, the test-set answers, which means you can score your own model locally instead of depending on a hosted leaderboard. The MMMU-Pro companion adds harder multiple-choice items and vision-only settings for a stricter read on genuine understanding. For teams comparing open-weight vision-language models, this is a practical, reproducible yardstick.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category