#198 · Primary category: MLOps & Evaluation
MultiBench
[NeurIPS 2021] Multiscale Benchmarks for Multimodal Representation Learning
Project last updated:01/27/24
GitHub Stars
636
Forks
100
Contributors
17
License
MIT
Why we included this project
Comparing multimodal models is usually a mess of incompatible setups, and MultiBench exists to fix that. It bundles 15 datasets spanning affective computing, healthcare, robotics, finance, HCI, and multimedia, and evaluates models on more than accuracy: it also tracks training and inference cost and how well a model holds up when modalities are noisy or missing. The companion MultiZoo toolkit ships roughly 20 fusion methods, objective functions, and training structures in modular form, so you can swap components instead of reimplementing every baseline. For researchers and teams that want reproducible comparisons, that combination is genuinely useful, and the documented process for adding new datasets and algorithms makes it easy to extend the benchmark to your own problem.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models