MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement Learning

fanqing meng, Lingxiao Du, kkkai (SII), Zhixiang Zhou, Quanfeng Lu, Fu, Botian Shi, wenhai.wang, JunjunHe, KAIPENG ZHANG, Ping Luo, Yu Qiao, Qiaosheng ZHANG, SII - Wenqi Shao

MM-Eureka: Exploring Visual Aha Moment with Rule-based Large-scale Reinforcement Learning: 61 upvotes on Hugging Face Daily Papers, #3 of 50 papers on 2025-03-11. Day-by-day upvote history.

We present MM-Eureka, a multimodal reasoning model that successfully extends large-scale rule-based reinforcement learning (RL) to multimodal reasoning. While rule-based RL has shown remarkable success in improving LLMs' reasoning abilities in text domains, its application to multimodal settings has remained challenging. Our work reproduces key characteristics of text-based RL systems like DeepSeek-R1 in the multimodal space, including steady increases in accuracy reward and response length, and the emergence of reflection behaviors. We demonstrate that both instruction-tuned and pre-trained models can develop strong multimodal reasoning capabilities through rule-based RL without supervised fine-tuning, showing superior data efficiency compared to alternative approaches. We open-source our complete pipeline to foster further research in this area. We release all our codes, models, data, etc. at https://github.com/ModalMinds/MM-EUREKA

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.