Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:17.658024Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 227 outbound references and 8 inbound Pith citation observations for arXiv:2505.18536.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:17.658024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T14:14:45.547957Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:50:10.284267Z
100 of 227 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 40097270-201b-4671-98a9-96a86ad89f99 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement learning: An introduction, volume 1
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb105126-c838-4241-9d62-ac96dbab593a · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Proximal Policy Optimization Algorithms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92fa4f2c-49c8-4842-beb4-81fb763b8b74 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Training language models to follow instructions with human feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc18dca8-c084-486a-bde8-b13d88490781 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69da4cf-d149-4c1b-96e6-1f6d07b198ca · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae63e3d-924b-4ddf-9049-9c858c37329e · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b9a938-15ac-413e-89a0-9c18cd2a4228 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o1
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 775fcef7-b5f1-4507-a9ea-236e8404c9b1 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9572b1d-50da-4296-9599-ef43d6891972 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mind with Eyes: from Language Reasoning to Multimodal Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c33b26-e119-48be-820e-93c2b9a5bca1 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1649ca34-a81d-4b4d-9a27-1477797c16f6 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A markovian decision process
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c754e2d2-b518-4e9b-a3cc-22f8c68a7543 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Convergence of q-learning: A simple proof
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5172b4c-9916-4ecd-89e9-3b917ef718a9 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Playing Atari with Deep Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94e0d25-b8ca-41a3-a2af-4be34aa47d87 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Human-level control through deep reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245fef20-d355-49fc-8b71-b4eb0183c79e · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Deep reinforcement learning with double q-learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4e90db-6a84-411d-acba-9da6c3092262 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dueling network architectures for deep reinforcement learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 549c1f6f-022e-4745-9135-a4973e29a50d · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Rainbow: Combining improvements in deep reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7111c2-ba27-468a-8ef9-5abf2e9bb760 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Policy gradi- ent methods for reinforcement learning with function approximation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9cc4bfa-6b53-4b20-84ba-52b228655f65 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Actor-critic algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e535d775-dcce-4472-9b7f-d40e5c2b0f2e · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Asynchronous methods for deep reinforce- ment learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97eef202-39fd-45c0-b930-3f4fc3ddc408 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Trust region policy optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44ec016-602a-47a9-9e9b-e07d14cb0a49 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a603a5-fbfa-4f22-9856-bbadb950af4b · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reasoning with large language models, a survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8130aecb-d762-4666-9385-a6c66663eb02 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2507e31d-74d1-4d0c-9646-f009e4797694 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d0d5d5-e971-4419-b65e-6fb3f6c96c97 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Thinking Machines: A Survey of LLM based Reasoning Strategies
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be25c8b-00c7-4257-a608-0a4dfca4605f · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b48aa1-d17b-4230-8db2-f10c5ecc337c · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2629d68-8aff-47db-b431-dec3ae56af11 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7babe830-f6d3-4d4c-853a-38d3ad66ec9d · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b47eff1-c366-4880-b0c7-63af299acbcb · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception, reason, think, and plan: A survey on large multimodal reasoning models, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1580e653-fd74-486c-9a03-a89e1b2efe30 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c63fa0b-ff1e-47dc-ba9f-59a724e5c016 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o3 and o4-mini
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16623895-9b4e-4516-b4a9-2942d29a3356 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Grok 3 beta — the age of reasoning agents
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86bcda32-3c84-4331-a4b3-9b6ad94135f3 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a962b26-e270-435c-9054-635f65b65d34 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec5c169-3d82-4565-a809-74ad2bb3a543 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3deeddac-e41a-4e02-8e2a-6e3b1a0ad8f3 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Approximating kl divergence
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a22cf9-07fd-4372-a117-a775ac9d9bc1 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d230f4-4c81-402d-850b-fcd968b02023 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bcbb80-eed5-449a-9c08-cdb14c4ad7e0 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa8e1e95-c984-4d5d-9d37-e36825a8a940 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Tinyzero
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061d8941-df23-4968-95e9-82c3a7d639e5 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ff927c-8585-4733-abf8-33827210c667 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051c8f17-2752-4dd2-86b8-fe78be8377d2 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65bc45ee-a926-45a0-a762-ba44de584247 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Audio-reasoner: Improving reasoning capability in large audio language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b9ad463-7fbf-40ee-8a5f-8390c59237f0 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da64ec5-3e68-4166-89e9-1a7e253e5cce · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e33bb8d-c9fa-47f6-aa5d-e271e2591963 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8cd184-7364-4401-a920-23f1239dae01 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1d1cfb-edff-4ec0-b4f5-2733b3adcbfc · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df3377d1-a140-4ef4-9cf3-330f7444868f · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536638f2-8005-48a4-aa21-835ba8a79f81 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8fc103-ae94-46b0-a2ee-6afb7aa41df0 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bba1d4a-7bf3-4a4a-86c7-0cf2c789b572 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vagen: Training vlm agents with multi-turn reinforcement learning, 2025
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b6a6df-454f-473b-a4f1-bc38cbc032f4 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e804e79d-632a-4614-aa32-b12fa31e5490 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Measuring multimodal mathematical reasoning with math-vision dataset
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ec617e7-3684-4a89-8b41-a39d5fd66cae · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081cfcc6-85eb-496c-9b7f-2c9ed819fd1c · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0cb03e-51e6-43a9-81be-841156b56272 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5df7d0-5c0f-4aeb-b61c-29c32bf1862e · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04b69c3-4337-40da-b1e5-a0d8569631fb · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676d8967-aa9c-42ae-b9b5-db8c76bef781 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a65190-299f-4f0e-b539-ec674f8f4100 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e892892b-c10d-45a6-8065-9432d3feab2f · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72cdbf10-5872-413f-8e80-172e4bc18ba5 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ecc574c-ada4-42b0-9d5d-6131135f8e9a · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e676ee6-46e1-4876-99b2-0bbc10440dfa · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa25072-aa78-4ace-88dd-3202fa3459b9 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d184a4-a0f3-43e3-a7be-ea742246e1a9 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5aa028-7c61-4799-ba77-a4ed484c335f · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afac0b4d-6a62-47bd-83f1-497cc0567c8d · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mmr1: Ad- vancing the frontiers of multimodal reasoning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f12bbf5-1ba4-436f-8aad-46720d7a2ad2 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be1b0c93-c556-486c-8a10-b029c698a2e9 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a457b09b-f3bc-4a6b-acf1-bffc6f8216c3 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ce94ea-7263-4968-aee8-9c4a75542481 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab00f3a-ea59-4597-8b1d-393e34eee0cd · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b345cae-a46b-427a-bb86-0ff66b6bf849 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Noisyrollout: Reinforcing visual reasoning with data augmentation
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00eac3aa-c4fa-40f9-ac6b-2a9297bbe228 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7a18cf-e7de-4faa-8a78-5e79f51dd4f3 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Fast-slow thinking for large vision-language model reasoning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1036505-1980-4395-87a2-f3dfba3bb0fd · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56723f23-61f3-4feb-a363-1b8e79caf6f7 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570fc3a6-5555-469b-82cf-12e948096f05 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a87c45-00b5-4cc3-8c36-faf3d4d6f4e7 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eff6400-1167-46f0-8655-a051a7ff64cf · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496011d8-4e0b-4c11-b078-090c603f2b73 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ee01f6-ba4b-4c0a-bea4-e51e7d7f9665 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcfc2b59-51ec-41ed-b2ca-a510e7d13a85 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Compile Scene Graphs with Reinforcement Learning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe23011-cb10-4f82-b0f2-5d19a96a19ce · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eee2c39-18e0-4c37-8fa5-68c866112ea6 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-track: Direct application of mllms to visual object tracking via reinforcement learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf62a27-321c-471b-b20c-074f2896c833 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SeekWorld: Geolocation is a natural RL task for o3-like visual clue-tracking
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8533d11d-b34c-4e6c-9ea3-fe6cdb35c3f8 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86b18bf-b91f-452c-bf36-547c9fa8a0c5 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6ac10c-57da-444c-8232-5b6303355bab · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61f2bfb-3c1d-4de3-9c5c-a0b9edd41418 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reason-rft: Reinforcement fine-tuning for visual reasoning
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967f09ed-0582-4c9f-8691-c1c707ee491e · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5e1ace-dbc9-410d-9784-e1158d796821 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea82e587-f15f-457d-b9f4-c78629abcb86 · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi-VL Technical Report
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5214968e-aa2d-443b-8d4e-ed8998ede8dd · outbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-vision: Let’s first take a look at the image
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9161386b-c4f9-478d-afcf-2bc15d5c4f0c · inbound
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c9dc37b-4062-4914-bf2e-5acfc1754921 · inbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc3102e1-4c8c-419d-be8b-9b6a5b0169e9 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7fac40b1-5516-4982-9ade-633169de5a16 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df4417a6-4e26-4106-8a73-5ae38a3056af · inbound
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42e34c38-0cf8-4028-9238-765ca4e6a02b · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 950009ec-c769-48be-929b-79169082bc1a · inbound
Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91d4c9ac-a3cf-4d38-9d23-bef70906eb8f · inbound
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.