#114 · Primary category: MLOps & Evaluation
VBench
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
Project last updated:08/21/26
GitHub Stars
1.8K
Forks
133
Contributors
21
License
Apache-2.0
Why we included this project
If you build or maintain text-to-video models, the hardest part is often telling whether a fresh version is genuinely better or just different. VBench addresses that by defining "video generation quality" as 16 separate dimensions, from subject consistency and motion smoothness to temporal flickering and spatial relationships, each with its own prompt suite and automated scoring pipeline. Because the benchmark was validated against human preference annotations across every dimension, its scores tend to line up with how real viewers perceive the output, which matters when you are comparing model candidates or chasing regressions. The repository ships the full prompt sets, generated sample videos, and unified evaluation code, and it extends to VBench-2.0 for broader capability coverage. For head-to-head comparisons, debugging a specific failure mode, or backing up a paper or product evaluation with credible numbers, this is a practical toolkit rather than a research toy.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models