#657 · Primary category: Education & Research

VisualThinker-R1-Zero

deepseek deepseek-r1 deepseek-r1-zero grpo multimodal multimodal-journey multimodal-r1 post-training r1 r1-zero reasoning reinforcement-learning

Explore the Multimodal “Aha Moment” on 2B Model

Project last updated:03/18/25

GitHub Stars

623

Forks

23

Contributors

3

License

Other

Why we included this project

This repo documents a research replication that applies DeepSeek-R1-Zero's reinforcement-learning recipe to a small vision-language model. The authors trained a 2B Qwen2-VL model with GRPO and no supervised fine-tuning, and report that it develops the same self-reflection and lengthened responses seen in the text-only R1-Zero runs, including the kind of 'aha moment' where the model catches and corrects its own mistakes. That makes it a useful reference for researchers and students working on multimodal reasoning, especially anyone curious about how reasoning behavior emerges during RL post-training. A checkpoint is released on Hugging Face, so it also works as a starting point for reproducing or extending the training setup. It is a research artifact rather than a production tool.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category