{"work":{"id":"d36889cb-edb6-448f-9a50-36df8b1623e5","openalex_id":null,"doi":null,"arxiv_id":"2504.07615","raw_key":null,"title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","authors":null,"authors_text":"Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao","year":2025,"venue":"cs.CV","abstract":"Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Motivated by this observation, we investigate the extension of R1-style reinforcement learning to Vision-Language Models (VLMs), aiming to enhance their visual reasoning capabilities. To this end, we develop VLM-R1, a dedicated framework designed to harness RL for improving VLMs' performance on general vision-language tasks. Using this framework, we further explore the feasibility of applying RL to visual domain. Experimental results indicate that the RL-based model not only delivers competitive performance on visual understanding tasks but also surpasses Supervised Fine-Tuning (SFT) in generalization ability. Furthermore, we conduct comprehensive ablation studies that uncover a series of noteworthy insights, including the presence of reward hacking in object detection, the emergence of the \"OD aha moment\", the impact of training data quality, and the scaling behavior of RL across different model sizes. Through these analyses, we aim to deepen the understanding of how reinforcement learning enhances the capabilities of vision-language models, and we hope our findings and open-source contributions will support continued progress in the vision-language RL community. Our code and model are available at https://github.com/om-ai-lab/VLM-R1","external_url":"https://arxiv.org/abs/2504.07615","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T03:17:51.904109+00:00","pith_arxiv_id":"2504.07615","created_at":"2026-05-09T06:15:38.911586+00:00","updated_at":"2026-07-11T03:17:51.904109+00:00","title_quality_ok":true,"display_title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","render_title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model"},"hub":{"state":{"work_id":"d36889cb-edb6-448f-9a50-36df8b1623e5","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":110,"external_cited_by_count":null,"distinct_field_count":8,"first_pith_cited_at":"2025-03-21T17:52:43+00:00","last_pith_cited_at":"2026-07-07T01:00:51+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T16:49:32.242000+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":15},{"context_role":"baseline","n":2},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":15},{"context_polarity":"baseline","n":2},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","claims":[{"claim_text":"Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Mot","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"from real competitions, spanning 16 mathematical dis- ciplines and five difficulty levels, each embedded in a visual context (figures, diagrams, plots). 8.ChartQA[27]: contains 9.6K human-written and 23.1K generated questions over diverse chart types, requiring both visual parsing and table/logic operations. Baselines.We compare our methods with multiple reasoning-oriented VLMs. 1.VLM-R1[34]: extends R1-style RLVR to VLMs by leveraging tasks with deterministic visual ground truth. 2.LMM-R1[30]:l","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"the-art MLLMs, comprising 3 proprietary models (GPT-5.2 [32], Gemini-3-flash- preview-nothinking[7],Grok-4-1-fast-no-reasoning[45])and12open-sourcemod- els (Kimi-VL-A3B-Instruct [36], Gemma3-4B [35], Qwen3-VL-4B-Instruct [47], Qwen3-VL-8B-Instruct [47], Qwen2.5-VL-3B-Instruct [2], InternVL3_5-8B [41], InternVL3_5-14B [41], QianFan-VL-8B [8], VLM-R1 [34], MiniCPM-V-4_5 [51], Deepseek_VL_7B [27], Deepseek-VL2-tiny [44]). For the proprietary models, we 10 Han et al. utilize APIs for testing. For op","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 4","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"offers a simple starting workflow for studying how agentic search can identify the right entity and bind it to the right visual instance. References [1] Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025. [2] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqia","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"groupwise RL method that enforces view-consistent and transformation-invariant objectives. Visionary-R1 [225] enforces image captioning as a prerequisite step before reasoning, mitigating shortcut exploitation during reinforcement finetuning. A line of curriculum-learning methods have also been proposed to ease and smooth the RL training process of vision reinforcement finetuning [226, 227, 228, 229, 217]. R1-V [227] introducesVLM-GymandtrainsG0/G1modelsviascalable,pureRLself-evolutionwithaperce","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"1 Variance of Importance Weights Equalsχ 2-Divergence Proposition 1.LetPandQbe two probability distributions withQabsolutely continuous with respect toP, and letρ(x) =P(x)/Q(x)denote the importance ratio. Then VarQ[ρ] =χ 2(P∥Q),(17) whereχ 2(P∥Q) =E Q \u0002 (ρ−1) 2\u0003 is theχ 2-divergence. Proof.By the definition of variance: VarQ[ρ] =E Q[ρ2]−(E Q[ρ])2 .(18) We first evaluate the mean ofρunderQ: EQ[ρ] = Z Q(x)· P(x) Q(x) dx= Z P(x)dx= 1,(19) where the last equality holds becausePis a valid probability","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-03T08:23:44.471966+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"95ea55a3-99e7-4d59-b52d-8331184ceda8","orcid":null,"display_name":"Haozhan Shen"},{"id":"f8915b26-557c-468c-aa44-7162a1a4de48","orcid":null,"display_name":"Peng Liu"},{"id":"facab044-1d99-4b08-940d-c8e9399432dd","orcid":null,"display_name":"Jingcheng Li"},{"id":"1cb69036-3d03-4e0b-8e35-6515db0de586","orcid":null,"display_name":"Chunxin Fang"},{"id":"893c15a6-6be4-4e2e-ac43-c2583f6949ee","orcid":null,"display_name":"Yibo Ma"},{"id":"fc7c2587-bc82-4d04-ab0b-dd82b904f481","orcid":null,"display_name":"Jiajia Liao"}]},"error":null,"updated_at":"2026-07-03T08:23:44.050528+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-17T05:49:42.865689+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":27},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":26},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":25},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":15},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":15},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":14},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":13},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":13},{"title":"Video-R1: Reinforcing Video Reasoning in MLLMs","work_id":"0ce88332-564c-4361-8e2a-3850eb1ace9c","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":11},{"title":"Visual-RFT: Visual Reinforcement Fine-Tuning","work_id":"872f09b5-998d-4a66-9a2f-f7ec2407cd62","shared_citers":10},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":9},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":9},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":9},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":9},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":8},{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","work_id":"5f6cf57b-2407-4127-b39c-d8a61494e474","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Perception-r1: Pioneering perception policy with reinforcement learning","work_id":"35656592-ffc7-4aef-9baf-0f8694c0c987","shared_citers":7},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":7},{"title":"R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization","work_id":"e1614961-16d2-43bc-908a-8c57da5b151c","shared_citers":7},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Kimi-VL Technical Report","work_id":"c876520f-8a20-44f3-b92a-bf7d35bd430f","shared_citers":6}],"time_series":[{"n":9,"year":2025},{"n":38,"year":2026}],"dependency_candidates":[{"n":1,"role":"baseline","polarity":"baseline","paper_title":"CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding","primary_cat":"cs.CV","context_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 41.03 46.23 7∼8B Models Mantis-8B [19] - 1.52 0.00 3.31 12.18 2.08 1.00 1.01 10.02 0.00 0.85 3.20 LLaVA-OV-7B [22] - 6.06 3.19 3.43 0.18 1.04 1.08 9.09 15.43 6.93 0.85 4.73 MiniCPM2.6-8B [45] - 14.58 2.13 14.","citing_arxiv_id":"2604.22498"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems","primary_cat":"cs.RO","context_text":"the failure feedback, enabling the planner to reason over the full planning history when generating the remaining plan. Table 11 below presents an example reasoning trace. Input: <observation> Object positions: Object 0: [0.75, 1.75, 0.08] . . . Object 1: [2.25, 1.75, 0.08] . . . Robot states: Robot 0: base=[0.25, 0.75, 0.00], arm=[0.35, 0.85, 0.10] . . . Robot 1: base=[2.75, 0.75, 0.00], arm=[2.65, 0.85, 0.10] . . . </observation> FULLPLANPlanner Reasoning: <think>Okay, let me analyze the given environment before coming up with a multi-step movement plan . . . I will make sure the joint plan is feasible for all robots under this configuration . . . Let me finalize this plan: it coordinates robots in parallel and avoids unnecessary motion . . . <think> [{\"Robot 0\": \"Move [0.","citing_arxiv_id":"2604.21138"}]},"error":null,"updated_at":"2026-05-17T05:49:38.628247+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-17T05:49:38.570595+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","claims":[{"claim_text":"Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Mot","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"from real competitions, spanning 16 mathematical dis- ciplines and five difficulty levels, each embedded in a visual context (figures, diagrams, plots). 8.ChartQA[27]: contains 9.6K human-written and 23.1K generated questions over diverse chart types, requiring both visual parsing and table/logic operations. Baselines.We compare our methods with multiple reasoning-oriented VLMs. 1.VLM-R1[34]: extends R1-style RLVR to VLMs by leveraging tasks with deterministic visual ground truth. 2.LMM-R1[30]:l","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"the-art MLLMs, comprising 3 proprietary models (GPT-5.2 [32], Gemini-3-flash- preview-nothinking[7],Grok-4-1-fast-no-reasoning[45])and12open-sourcemod- els (Kimi-VL-A3B-Instruct [36], Gemma3-4B [35], Qwen3-VL-4B-Instruct [47], Qwen3-VL-8B-Instruct [47], Qwen2.5-VL-3B-Instruct [2], InternVL3_5-8B [41], InternVL3_5-14B [41], QianFan-VL-8B [8], VLM-R1 [34], MiniCPM-V-4_5 [51], Deepseek_VL_7B [27], Deepseek-VL2-tiny [44]). For the proprietary models, we 10 Han et al. utilize APIs for testing. For op","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 4","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"offers a simple starting workflow for studying how agentic search can identify the right entity and bind it to the right visual instance. References [1] Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025. [2] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqia","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"groupwise RL method that enforces view-consistent and transformation-invariant objectives. Visionary-R1 [225] enforces image captioning as a prerequisite step before reasoning, mitigating shortcut exploitation during reinforcement finetuning. A line of curriculum-learning methods have also been proposed to ease and smooth the RL training process of vision reinforcement finetuning [226, 227, 228, 229, 217]. R1-V [227] introducesVLM-GymandtrainsG0/G1modelsviascalable,pureRLself-evolutionwithaperce","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"1 Variance of Importance Weights Equalsχ 2-Divergence Proposition 1.LetPandQbe two probability distributions withQabsolutely continuous with respect toP, and letρ(x) =P(x)/Q(x)denote the importance ratio. Then VarQ[ρ] =χ 2(P∥Q),(17) whereχ 2(P∥Q) =E Q \u0002 (ρ−1) 2\u0003 is theχ 2-divergence. Proof.By the definition of variance: VarQ[ρ] =E Q[ρ2]−(E Q[ρ])2 .(18) We first evaluate the mean ofρunderQ: EQ[ρ] = Z Q(x)· P(x) Q(x) dx= Z P(x)dx= 1,(19) where the last equality holds becausePis a valid probability","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-03T08:23:44.467704+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","claims":[{"claim_text":"Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Mot","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 4","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"offers a simple starting workflow for studying how agentic search can identify the right entity and bind it to the right visual instance. References [1] Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025. [2] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqia","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"your large multimodal model achieve human-like mathemat- ical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20023-20070, 2025. 3, 4 [24] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Rad- ford, and Oleg Klimov. Proximal Policy Optimization Algo- rithms, 2017. 3 [25] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al. Vlm","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Ensembling multiple models or using rule-based environments like CAD compilers can further neutralize proxy flaws [223, 224]. Trajectory and Optimization Interventions exploit the iterative nature of generation. Timestep-aware schemes, such as temporal asymmetric interventions [171] or dynamic distortion-perception weighting [211], decay the proxy reward's influence over time to preserve the global structure. Directional shaping, such as D2-Align [193], cor- rects the optimization direction in e","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"[30] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.CoRR, abs/1707.06347, 2017. [31] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.CoRR, abs/2402.03300, 2024. [32] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Q","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"[23] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017. [24] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024. [25] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (5 contexts).","role_counts":[{"n":5,"context_role":"background"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-17T05:49:42.869731+00:00"}},"summary":{"title":"VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model","claims":[{"claim_text":"Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Mot","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 4","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"offers a simple starting workflow for studying how agentic search can identify the right entity and bind it to the right visual instance. References [1] Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025. [2] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqia","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"your large multimodal model achieve human-like mathemat- ical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20023-20070, 2025. 3, 4 [24] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Rad- ford, and Oleg Klimov. Proximal Policy Optimization Algo- rithms, 2017. 3 [25] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al. Vlm","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Ensembling multiple models or using rule-based environments like CAD compilers can further neutralize proxy flaws [223, 224]. Trajectory and Optimization Interventions exploit the iterative nature of generation. Timestep-aware schemes, such as temporal asymmetric interventions [171] or dynamic distortion-perception weighting [211], decay the proxy reward's influence over time to preserve the global structure. Directional shaping, such as D2-Align [193], cor- rects the optimization direction in e","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"[30] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.CoRR, abs/1707.06347, 2017. [31] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.CoRR, abs/2402.03300, 2024. [32] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Q","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"[23] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017. [24] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024. [25] Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (5 contexts).","role_counts":[{"n":5,"context_role":"background"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"graph":{"co_cited":[{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":27},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":26},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":25},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":15},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":15},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":14},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":13},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":13},{"title":"Video-R1: Reinforcing Video Reasoning in MLLMs","work_id":"0ce88332-564c-4361-8e2a-3850eb1ace9c","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":11},{"title":"Visual-RFT: Visual Reinforcement Fine-Tuning","work_id":"872f09b5-998d-4a66-9a2f-f7ec2407cd62","shared_citers":10},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":9},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":9},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":9},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":9},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":8},{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","work_id":"5f6cf57b-2407-4127-b39c-d8a61494e474","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Perception-r1: Pioneering perception policy with reinforcement learning","work_id":"35656592-ffc7-4aef-9baf-0f8694c0c987","shared_citers":7},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":7},{"title":"R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization","work_id":"e1614961-16d2-43bc-908a-8c57da5b151c","shared_citers":7},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Kimi-VL Technical Report","work_id":"c876520f-8a20-44f3-b92a-bf7d35bd430f","shared_citers":6}],"time_series":[{"n":9,"year":2025},{"n":38,"year":2026}],"dependency_candidates":[{"n":1,"role":"baseline","polarity":"baseline","paper_title":"CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding","primary_cat":"cs.CV","context_text":"87 98.00 100.00 96.88 100.00 98.99 91.06 92.08 97.44 97.18 Large-Scale Models (≥70B) LLaVA-OV-72B [22] - 13.26 5.34 26.84 12.91 7.64 2.14 17.83 21.60 11.88 8.55 13.65 InternVL2-76B [40] - 15.91 10.64 36.40 30.73 20.83 5.74 46.46 41.28 32.67 26.50 26.72 InternVL3-78B [61] - 10.04 9.57 24.12 27.08 14.58 10.44 50.51 38.08 45.54 17.09 24.71 Qwen2-VL-72B [42] - 46.12 46.81 64.46 26.73 22.57 18.62 33.33 62.53 50.50 17.09 38.88 Qwen2.5-VL-72B [3] - 43.75 46.81 69.98 34.32 29.17 8.31 62.63 59.92 66.34 41.03 46.23 7∼8B Models Mantis-8B [19] - 1.52 0.00 3.31 12.18 2.08 1.00 1.01 10.02 0.00 0.85 3.20 LLaVA-OV-7B [22] - 6.06 3.19 3.43 0.18 1.04 1.08 9.09 15.43 6.93 0.85 4.73 MiniCPM2.6-8B [45] - 14.58 2.13 14.","citing_arxiv_id":"2604.22498"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems","primary_cat":"cs.RO","context_text":"the failure feedback, enabling the planner to reason over the full planning history when generating the remaining plan. Table 11 below presents an example reasoning trace. Input: <observation> Object positions: Object 0: [0.75, 1.75, 0.08] . . . Object 1: [2.25, 1.75, 0.08] . . . Robot states: Robot 0: base=[0.25, 0.75, 0.00], arm=[0.35, 0.85, 0.10] . . . Robot 1: base=[2.75, 0.75, 0.00], arm=[2.65, 0.85, 0.10] . . . </observation> FULLPLANPlanner Reasoning: <think>Okay, let me analyze the given environment before coming up with a multi-step movement plan . . . I will make sure the joint plan is feasible for all robots under this configuration . . . Let me finalize this plan: it coordinates robots in parallel and avoids unnecessary motion . . . <think> [{\"Robot 0\": \"Move [0.","citing_arxiv_id":"2604.21138"}]},"authors":[{"id":"1cb69036-3d03-4e0b-8e35-6515db0de586","orcid":null,"display_name":"Chunxin Fang","source":"manual","import_confidence":0.72},{"id":"95ea55a3-99e7-4d59-b52d-8331184ceda8","orcid":null,"display_name":"Haozhan Shen","source":"manual","import_confidence":0.72},{"id":"fc7c2587-bc82-4d04-ab0b-dd82b904f481","orcid":null,"display_name":"Jiajia Liao","source":"manual","import_confidence":0.72},{"id":"facab044-1d99-4b08-940d-c8e9399432dd","orcid":null,"display_name":"Jingcheng Li","source":"manual","import_confidence":0.72},{"id":"f8915b26-557c-468c-aa44-7162a1a4de48","orcid":null,"display_name":"Peng Liu","source":"manual","import_confidence":0.72},{"id":"893c15a6-6be4-4e2e-ac43-c2583f6949ee","orcid":null,"display_name":"Yibo Ma","source":"manual","import_confidence":0.72}]}}