{"work":{"id":"b13ff087-9614-48cc-8991-ad75b6543bbc","openalex_id":"https://openalex.org/W4415153693","doi":"10.48550/arxiv.2504.08837","arxiv_id":"2504.08837","raw_key":null,"title":"VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning","authors":null,"authors_text":"Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, Wenhu Chen","year":2025,"venue":"cs.LG","abstract":"Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such as GPT-4o, on various math and science benchmarks. However, their multimodal reasoning capabilities remain on par with fast-thinking models. For instance, GPT-o1's performance on benchmarks like MathVista, MathVerse, and MathVision is similar to fast-thinking models. In this paper, we aim to enhance the slow-thinking capabilities of vision-language models using reinforcement learning (without relying on distillation) to advance the state of the art. First, we adapt the GRPO algorithm with a novel technique called Selective Sample Replay (SSR) to address the vanishing advantages problem. While this approach yields strong performance, the resulting RL-trained models exhibit limited self-reflection or self-verification. To further encourage slow-thinking, we introduce Forced Rethinking, which appends a rethinking trigger token to the end of rollouts in RL training, explicitly enforcing a self-reflection reasoning step. By combining these two techniques, our model, VL-Rethinker, advances state-of-the-art scores on MathVista, MathVerse to achieve 80.4%, 63.5% respectively. VL-Rethinker also achieves open-source SoTA on multi-disciplinary benchmarks such as MathVision, MMMU-Pro, EMMA, and MEGA-Bench, narrowing the gap with OpenAI-o1. Our empirical results show the effectiveness of our approaches.","external_url":"https://arxiv.org/abs/2504.08837","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2504.08837","created_at":"2026-05-10T11:58:58.816830+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning","render_title":"VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning"},"hub":{"state":{"work_id":"b13ff087-9614-48cc-8991-ad75b6543bbc","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":64,"external_cited_by_count":0,"distinct_field_count":6,"first_pith_cited_at":"2025-03-21T17:52:43+00:00","last_pith_cited_at":"2026-07-07T03:50:43+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T10:49:29.317852+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":8},{"context_role":"baseline","n":5},{"context_role":"dataset","n":3}],"polarity_counts":[{"context_polarity":"background","n":8},{"context_polarity":"baseline","n":5},{"context_polarity":"use_dataset","n":3}],"runs":{},"summary":{},"graph":{},"authors":[]}}