Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 100 of 227 outbound references and 8 inbound Pith citation observations for arXiv:2505.18536.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18536 v1

Coverage vector

measured 100 of 227 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:17.658024Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T14:14:45.547957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:10.284267Z

Reference resolution

100 of 227 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40097270-201b-4671-98a9-96a86ad89f99 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement learning: An introduction, volume 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.011412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.011412Z digest=sha256:730f064f41cbaacbc526f9a41d40317156bbb0c4c77610fee97a15cee93ddbd2

Observation cb105126-c838-4241-9d62-ac96dbab593a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Proximal Policy Optimization Algorithms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.103283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.103283Z digest=sha256:c59813e4836396e4675d5b3d7b5267c32638b2073b261d1c22bd690771a7bc02

Observation 92fa4f2c-49c8-4842-beb4-81fb763b8b74 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Training language models to follow instructions with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.256788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.256788Z digest=sha256:9f426201c42acb046e3dbf0a6b855fe30ebf72bffbbadae3a77f6f7ae677e78b

Observation dc18dca8-c084-486a-bde8-b13d88490781 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.367729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.367729Z digest=sha256:ad60e72386da3cc6668a3b7979b01ccb7d875c6c8d58039c9d01e18135951473

Observation b69da4cf-d149-4c1b-96e6-1f6d07b198ca · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.535029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.535029Z digest=sha256:aff16785285debd530df22a1414f2b452e5d0b1207413a71d707c4843df7ba49

Observation 3ae63e3d-924b-4ddf-9049-9c858c37329e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.639365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.639365Z digest=sha256:37c42a05cac63aa3ccb2c7b699546da9ff83b40307f704e542d8d6f283860a79

Observation 81b9a938-15ac-413e-89a0-9c18cd2a4228 · outbound

This paper cites Introducing openai o1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o1

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.792427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.792427Z digest=sha256:6a1a267d22517095f5da813ddaada91ada60c2cefd091582288c6b4ec1b8fd18

Observation 775fcef7-b5f1-4507-a9ea-236e8404c9b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.923108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.923108Z digest=sha256:2633a1171d40314b410241429571950dac5f6b7971b98bb71bebaee8f811c2fd

Observation e9572b1d-50da-4296-9599-ef43d6891972 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.013331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.013331Z digest=sha256:dcbe9a3a2a5bcbaf248543bc8431b7e7178bc5e947b2d1a89fc882108365ed88

Observation 05c33b26-e119-48be-820e-93c2b9a5bca1 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.127322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.127322Z digest=sha256:9fff00ba77231ea1876b35a70846acdcd3062675dc6604a0a20175338f72f898

Observation 1649ca34-a81d-4b4d-9a27-1477797c16f6 · outbound

This paper cites A markovian decision process.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A markovian decision process

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.235317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.235317Z digest=sha256:4a405f375a5b5d3a18f4613ae07ece41995cfc470c43f1c4fb1d0adbf4b50f60

Observation c754e2d2-b518-4e9b-a3cc-22f8c68a7543 · outbound

This paper cites Convergence of q-learning: A simple proof.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Convergence of q-learning: A simple proof

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.350959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.350959Z digest=sha256:da3c6a0556b15e26f8f711acf6fb835d54c5b54bc0c8c8d5479e9f98ef87aa97

Observation b5172b4c-9916-4ecd-89e9-3b917ef718a9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Playing Atari with Deep Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.496847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.496847Z digest=sha256:e431c88e71f88203f3f299e040f016c4156490d1e144863474ec532854296db4

Observation d94e0d25-b8ca-41a3-a2af-4be34aa47d87 · outbound

This paper cites Human-level control through deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Human-level control through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.619866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.619866Z digest=sha256:85a8d964ff0cf2eb7c6ef17d13ecc8b65b58735ef41da73639bc052fc1f60783

Observation 245fef20-d355-49fc-8b71-b4eb0183c79e · outbound

This paper cites Deep reinforcement learning with double q-learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Deep reinforcement learning with double q-learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.734580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.734580Z digest=sha256:1d47742406a158276dda047cf9b304465c0123e2264a0507eef2bdb765438c12

Observation 2a4e90db-6a84-411d-acba-9da6c3092262 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dueling network architectures for deep reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.823430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.823430Z digest=sha256:bb847a5eb6277bd316b081d632e0eabe4f4b6d9386fd87dbb7597e3c9e8a2ef3

Observation 549c1f6f-022e-4745-9135-a4973e29a50d · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Rainbow: Combining improvements in deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.948269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.948269Z digest=sha256:3e47dc859821bba39c8688c773a2de66c503962ee25376bf228c141bbc60566e

Observation cf7111c2-ba27-468a-8ef9-5abf2e9bb760 · outbound

This paper cites Policy gradi- ent methods for reinforcement learning with function approximation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Policy gradi- ent methods for reinforcement learning with function approximation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.081908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.081908Z digest=sha256:74c9075d52001d533f7668ed92a4cf4a00947a3cd878db813d07c86b89016f1f

Observation f9cc4bfa-6b53-4b20-84ba-52b228655f65 · outbound

This paper cites Actor-critic algorithms.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Actor-critic algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.143958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.143958Z digest=sha256:0d4e2b1afc16377eda30240be1bae70d817dd4b4db83d38c5c57fa851fae6691

Observation e535d775-dcce-4472-9b7f-d40e5c2b0f2e · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Asynchronous methods for deep reinforce- ment learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.255173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.255173Z digest=sha256:fd1cd0789c578191a4f5405169a4c35f0efc1842d6caf15fe39e7c3f163c4abd

Observation 97eef202-39fd-45c0-b930-3f4fc3ddc408 · outbound

This paper cites Trust region policy optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Trust region policy optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.401052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.401052Z digest=sha256:eb5a442425e12dfe0814badc858fe8e32b1f10e66ef078f41dd509f84263287a

Observation e44ec016-602a-47a9-9e9b-e07d14cb0a49 · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.524240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.524240Z digest=sha256:371ac108e8dfa7a8163f3b95d10fe170bfe47d2891044bc09e8e135a19d857cd

Observation 18a603a5-fbfa-4f22-9856-bbadb950af4b · outbound

This paper cites Reasoning with large language models, a survey.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reasoning with large language models, a survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.596496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.596496Z digest=sha256:214da5225540aeb04c83c33eabb710a5b1a329b9264327795d954511ba5246bb

Observation 8130aecb-d762-4666-9385-a6c66663eb02 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.735951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.735951Z digest=sha256:b301f4edb86f6956255b39efca8870577d3acf5858fe6be7708511f0747ee560

Observation 2507e31d-74d1-4d0c-9646-f009e4797694 · outbound

This paper cites A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.828360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.828360Z digest=sha256:fd50f55c180ecd463affd65697933ce461b55e849743a4cedb1e5574239727ed

Observation c7d0d5d5-e971-4419-b65e-6fb3f6c96c97 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.991423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.991423Z digest=sha256:cf112d1ddfeae515a29d78d66a7baee474eeadb8d9ad4efac9ff6e5cabfbbac6

Observation 4be25c8b-00c7-4257-a608-0a4dfca4605f · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.116827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.116827Z digest=sha256:774d766df69834efa343f7f1983d8acd4a14ad4b50e0706879e5842f156e5850

Observation 41b48aa1-d17b-4230-8db2-f10c5ecc337c · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.242194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.242194Z digest=sha256:989f74ab896053774562728dca5aeb020110d38936ffad537d8a7dbc8d206bad

Observation f2629d68-8aff-47db-b431-dec3ae56af11 · outbound

This paper cites Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.393683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.393683Z digest=sha256:797c10e106edd1d37ba53ebc65b6f5603edf34917892f1b19ee9ef8ae4e65574

Observation 7babe830-f6d3-4d4c-853a-38d3ad66ec9d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.547423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.547423Z digest=sha256:ee540ca760299bd91c4d3eebdf53d9c982cff4fb41b23037041f6395a3ea308a

Observation 7b47eff1-c366-4880-b0c7-63af299acbcb · outbound

This paper cites Perception, reason, think, and plan: A survey on large multimodal reasoning models, 2025.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception, reason, think, and plan: A survey on large multimodal reasoning models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.658758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.658758Z digest=sha256:4eb079258f40d5c4ddd2ef39d8b9d5f7059eaef0abdd8eb94db8fa09fd4649e6

Observation 1580e653-fd74-486c-9a03-a89e1b2efe30 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.745517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.745517Z digest=sha256:8c91baf371202073100d9fa9a8bd3e29beb1a5e36f9e368dcb4996d437582657

Observation 1c63fa0b-ff1e-47dc-ba9f-59a724e5c016 · outbound

This paper cites Introducing openai o3 and o4-mini.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o3 and o4-mini

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.857829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.857829Z digest=sha256:06d662f944164662d199949bc4dc9d495611989365509f87ea50ff7b0802651a

Observation 16623895-9b4e-4516-b4a9-2942d29a3356 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Grok 3 beta — the age of reasoning agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.968928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.968928Z digest=sha256:68a1465ffd11422688ffdc817692f5bf4eaaa75db579adfeb74751a21fdbeb9e

Observation 86bcda32-3c84-4331-a4b3-9b6ad94135f3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.068390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.068390Z digest=sha256:692daf71bef47a37f57ce9f04d962e4355e2340003bb08f584379a9a7de91ee7

Observation 5a962b26-e270-435c-9054-635f65b65d34 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.204895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.204895Z digest=sha256:9ecfdc354215f2178280115c6e5acae921f03f43e40b1ee5b2c2d6b794690e25

Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.361455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.361455Z digest=sha256:3c799b75ea382dae2887845b0fcbe1d37926c927b811b30346fa4e41761c3b18

Observation cec5c169-3d82-4565-a809-74ad2bb3a543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.476155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.476155Z digest=sha256:6429c8de52c8c3e6ec228a60eac0ed737df65a099569421f68c8469380834617

Observation 3deeddac-e41a-4e02-8e2a-6e3b1a0ad8f3 · outbound

This paper cites Approximating kl divergence.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Approximating kl divergence

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.595428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.595428Z digest=sha256:33f39ce7d8cd6816f0e81999b3a8f7e74c48be4529d1e1f9772963b4e51cf915

Observation d9a22cf9-07fd-4372-a117-a775ac9d9bc1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.705003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.705003Z digest=sha256:66cfa705923466dc8f276161bab51f7c611e184abeb9c195feacb5e3d2c236f0

Observation 19d230f4-4c81-402d-850b-fcd968b02023 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.807492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.807492Z digest=sha256:7d5499faa890c006251f3dcda79cf537d5c7f84230b98d24bc6605580722b531

Observation 58bcbb80-eed5-449a-9c08-cdb14c4ad7e0 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.891931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.891931Z digest=sha256:170747ab9888e59cde395c43868b1f85fc512ab6ac427d91e8251e08dceb380b

Observation fa8e1e95-c984-4d5d-9d37-e36825a8a940 · outbound

This paper cites Tinyzero.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Tinyzero

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.022145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.022145Z digest=sha256:0c42d64530e306b34a5e05630173270d231c3a91bc5848fd8042b25134ef9036

Observation 061d8941-df23-4968-95e9-82c3a7d639e5 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.154734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.154734Z digest=sha256:c30146d5a184ffcf7bbee7b0c0085274ad635d486a0c08197cf37b6e73766b1e

Observation 91ff927c-8585-4733-abf8-33827210c667 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.239675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.239675Z digest=sha256:10aa0b67e3ff88097f7d7e8e189f4e34ff89a1e7f922f127d40538ed4eb041b0

Observation 051c8f17-2752-4dd2-86b8-fe78be8377d2 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.310672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.310672Z digest=sha256:8c74b41bb0c8cc0a17989841689904a0dc6546270614fe625c08cb0cd4de5d04

Observation 65bc45ee-a926-45a0-a762-ba44de584247 · outbound

This paper cites Audio-reasoner: Improving reasoning capability in large audio language models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Audio-reasoner: Improving reasoning capability in large audio language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.413579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.413579Z digest=sha256:4e9ce41a7c6873339de8f81b4a8230709e85b1772d3688c3130ac4269c87dbfa

Observation 7b9ad463-7fbf-40ee-8a5f-8390c59237f0 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.461419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.461419Z digest=sha256:90b826ae242e2fd7239b8256f335f83580a16456cd42aaa0fe3b13debdcefa03

Observation 4da64ec5-3e68-4166-89e9-1a7e253e5cce · outbound

This paper cites SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.506592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.506592Z digest=sha256:2376aa94e154800a4a3c1193377cd2c6811e993c492ddea8b060adc0c2a4b1e4

Observation 0e33bb8d-c9fa-47f6-aa5d-e271e2591963 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.565978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.565978Z digest=sha256:4944a6b8455aa7336331a63570200e360951a8bd664e1d3d3c588233c29238ee

Observation bb8cd184-7364-4401-a920-23f1239dae01 · outbound

This paper cites EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.645079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.645079Z digest=sha256:1d0f336cc73ac1cdd0c44d72308154a123f6d5351df8df7d66e1c9f286430fbb

Observation 5c1d1cfb-edff-4ec0-b4f5-2733b3adcbfc · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.700122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.700122Z digest=sha256:d49836f50a41fae038a1918e11b63dfa7167ee18ec999281988280979e3191e6

Observation df3377d1-a140-4ef4-9cf3-330f7444868f · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.747182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.747182Z digest=sha256:192cc92587b7d10be162b924adfaae9d05dca9557abe8475757215b070973de4

Observation 536638f2-8005-48a4-aa21-835ba8a79f81 · outbound

This paper cites InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.797112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.797112Z digest=sha256:bb592fe5a6a37584d1e213dcf354330155cf1945d9526c6cda0fcda777c578ab

Observation 3f8fc103-ae94-46b0-a2ee-6afb7aa41df0 · outbound

This paper cites Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.851407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.851407Z digest=sha256:6d427c0f3cd488150d5d51bb3c5682ca3d12555cd30e434b14130c1c7e49b112

Observation 5bba1d4a-7bf3-4a4a-86c7-0cf2c789b572 · outbound

This paper cites Vagen: Training vlm agents with multi-turn reinforcement learning, 2025.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vagen: Training vlm agents with multi-turn reinforcement learning, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.898943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.898943Z digest=sha256:4954cd4d7d71f07ff8b27d8a4b67121a7a9034061ea1718b46b69e2f0f2cd5a7

Observation f9b6a6df-454f-473b-a4f1-bc38cbc032f4 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.950384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.950384Z digest=sha256:3d9e952827f57e327e9b34d9f7d2daee0c0e4f49c2eadac2e0f6c174b2d57c9f

Observation e804e79d-632a-4614-aa32-b12fa31e5490 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Measuring multimodal mathematical reasoning with math-vision dataset

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.994703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.994703Z digest=sha256:dac1aa5c3b7528d54efe09d652915d7ca46e7ad8945a1599ccee0866d5a32197

Observation 8ec617e7-3684-4a89-8b41-a39d5fd66cae · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.050530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.050530Z digest=sha256:1074ba96254e9560358b76c0db2e3d2dcdf0f60fb508a36a2d0c24b9ddca5125

Observation 081cfcc6-85eb-496c-9b7f-2c9ed819fd1c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.095636Z digest=sha256:bbc9c52b27288d64282f44e2efd9139c60157c4d88bd64e705558c7f5b0d54bf

Observation 7e0cb03e-51e6-43a9-81be-841156b56272 · outbound

This paper cites Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.148793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.148793Z digest=sha256:2509d26bb148a35234db6ae5b997ecaf2e2d38415dccaa3aaff0e33e43274f84

Observation cc5df7d0-5c0f-4aeb-b61c-29c32bf1862e · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.198563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.198563Z digest=sha256:cb1bb2bbd1b45b6203bc648afa4dc672abc304dfb3fd4908d3beb9e2085ec73d

Observation a04b69c3-4337-40da-b1e5-a0d8569631fb · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.248813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.248813Z digest=sha256:7eb9b434501354654cdb349a06a432febb05e5a8a92a3cc90ba3b1f0d9f70735

Observation 676d8967-aa9c-42ae-b9b5-db8c76bef781 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.353463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.353463Z digest=sha256:f2a36407df52deac9909998f7d8bd3bc47252bb73a7f33b7431957e03b8e06f1

Observation b2a65190-299f-4f0e-b539-ec674f8f4100 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.410946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.410946Z digest=sha256:f5b3c94121041d005cab340e63f4ee9ec543bf7fadbe3b634bf0f74aecdd4f06

Observation e892892b-c10d-45a6-8065-9432d3feab2f · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.468895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.468895Z digest=sha256:ace59aa4007a6c023b96d08e1e46ee01fb2c4f38ccb53e1ffb98a7b413302aac

Observation 72cdbf10-5872-413f-8e80-172e4bc18ba5 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.524790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.524790Z digest=sha256:76feda90cae0e3d953e676ad9c89e13d9ea163ef1581a889d8a28ca596db774c

Observation 9ecc574c-ada4-42b0-9d5d-6131135f8e9a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.568519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.568519Z digest=sha256:dc8d40d323f8186f448623044e252913f5b2fb4ff157ad4783c48d24eccdd30d

Observation 1e676ee6-46e1-4876-99b2-0bbc10440dfa · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.613909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.613909Z digest=sha256:4a59c503fc04c2c5d2734a422b7c8e0b4317edf8046df61dc184eb7d894c6b5c

Observation 9fa25072-aa78-4ace-88dd-3202fa3459b9 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.672507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.672507Z digest=sha256:51b171a74229495d8f0b2082f829d3c8c183aafa687927adad05938e0ab417f5

Observation 79d184a4-a0f3-43e3-a7be-ea742246e1a9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.730553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.730553Z digest=sha256:7650cdae8f8c1a09bd2128f5aa8b0c214bbc259738db5c142531367dee86030f

Observation 5f5aa028-7c61-4799-ba77-a4ed484c335f · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.783221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.783221Z digest=sha256:49e4f09ae017c21d8ed625c2bb0ddf0c6f7bb4fcbbca1163bec37572c9a50bca

Observation afac0b4d-6a62-47bd-83f1-497cc0567c8d · outbound

This paper cites Mmr1: Ad- vancing the frontiers of multimodal reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mmr1: Ad- vancing the frontiers of multimodal reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.842854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.842854Z digest=sha256:5b431679f423d4d607faee7dc3d357406efd9e69190a93a0049d7a81612c1481

Observation 1f12bbf5-1ba4-436f-8aad-46720d7a2ad2 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.922706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.922706Z digest=sha256:0f8b718667d3aae655eaff3fe4944e97a054f468bcf0af09599f1ca98af45c49

Observation be1b0c93-c556-486c-8a10-b029c698a2e9 · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.986136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.986136Z digest=sha256:e075ea0bc8c085305864d95bce858e507ea8742b0fffb5a88b637ab7e04f8cf4

Observation a457b09b-f3bc-4a6b-acf1-bffc6f8216c3 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.058067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.058067Z digest=sha256:b84d500a48205906aed7442296f256cc0fc2e314cdeb88519f0775598e4217b2

Observation e6ce94ea-7263-4968-aee8-9c4a75542481 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.106642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.106642Z digest=sha256:a43fff4bbd7cc6dafdfd35c3edcc262bd77f8c082295913424bbeac07182c671

Observation fab00f3a-ea59-4597-8b1d-393e34eee0cd · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.147711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.147711Z digest=sha256:679396c39578b7230212196d513716b55651221c56875954d48f54358fc493f4

Observation 5b345cae-a46b-427a-bb86-0ff66b6bf849 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.223794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.223794Z digest=sha256:1236d9f87beb1cb343b65b98c50d98d9e2496140d9b65b1f7c27b455089842d5

Observation 00eac3aa-c4fa-40f9-ac6b-2a9297bbe228 · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.321786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.321786Z digest=sha256:fc0d653f677ccab50937a678f149274217c3e15cc3b85ba5dbe1e25fd8ebaa62

Observation 4a7a18cf-e7de-4faa-8a78-5e79f51dd4f3 · outbound

This paper cites Fast-slow thinking for large vision-language model reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Fast-slow thinking for large vision-language model reasoning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.448124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.448124Z digest=sha256:c6f49b12886bf5dea1517d5aa9577a95a5f1d5afc708b94f674d2bdcadd264d9

Observation a1036505-1980-4395-87a2-f3dfba3bb0fd · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.565534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.565534Z digest=sha256:5dc2396b8ad4dd9cbd2640b2aae2dc11e8885dbf9bae45deaa4f8d1aadad2462

Observation 56723f23-61f3-4feb-a363-1b8e79caf6f7 · outbound

This paper cites Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.674647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.674647Z digest=sha256:e601f222e3c8f79b1f9e767a33120563c12ada9dca3f0115644a67e95abff2a8

Observation 570fc3a6-5555-469b-82cf-12e948096f05 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.750218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.750218Z digest=sha256:de4ed4685789a0f2bfe1ba16c7c0ed533afaadd03b2b84bd12c592056aff4409

Observation 21a87c45-00b5-4cc3-8c36-faf3d4d6f4e7 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.837024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.837024Z digest=sha256:7125b665d674694ce7ba4e04a582209367b18a159f4933aff05b167fd3cbacca

Observation 0eff6400-1167-46f0-8655-a051a7ff64cf · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.937329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.937329Z digest=sha256:6c0beb638635a527c51292db4501aa2ff20fa4f13af8ad9f2c1fcd3b8b45234b

Observation 496011d8-4e0b-4c11-b078-090c603f2b73 · outbound

This paper cites Q-Insight: Understanding Image Quality via Visual Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.030204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.030204Z digest=sha256:e05c8c01a542ab39a7996f3ed6b48056473a9319c2e268f39d367d1868dc85de

Observation 54ee01f6-ba4b-4c0a-bea4-e51e7d7f9665 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.116604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.116604Z digest=sha256:5103b8681b7537b0009c77065f8b325108c48a06e50a78ab6667ae1b3aac2fb2

Observation dcfc2b59-51ec-41ed-b2ca-a510e7d13a85 · outbound

This paper cites Compile Scene Graphs with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Compile Scene Graphs with Reinforcement Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.244853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.244853Z digest=sha256:853bf50b0d3ede84f669922e1a6504a5ad650d5a758a493cbed7fef30f40db0d

Observation 6fe23011-cb10-4f82-b0f2-5d19a96a19ce · outbound

This paper cites Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.408178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.408178Z digest=sha256:2361a6a8c4fea2a5e9baa53c71da61b9fb8d1e00efd6448958fc86eb6a9b5806

Observation 1eee2c39-18e0-4c37-8fa5-68c866112ea6 · outbound

This paper cites R1-track: Direct application of mllms to visual object tracking via reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-track: Direct application of mllms to visual object tracking via reinforcement learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.533191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.533191Z digest=sha256:5743c49993f54c7c36f4f874e247dce4dd06e11e54e399490a0c4c6bc3675081

Observation 5cf62a27-321c-471b-b20c-074f2896c833 · outbound

This paper cites SeekWorld: Geolocation is a natural RL task for o3-like visual clue-tracking.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SeekWorld: Geolocation is a natural RL task for o3-like visual clue-tracking

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.665298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.665298Z digest=sha256:e2d72b52876906449da56514a7afe4133a49062b171012a824046ce3b74a56d5

Observation 8533d11d-b34c-4e6c-9ea3-fe6cdb35c3f8 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.855930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.855930Z digest=sha256:bbcf0e6ad31d4431e18135077e2bd82538ff9be7102bdbe36146c59df5767739

Observation b86b18bf-b91f-452c-bf36-547c9fa8a0c5 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.974601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.974601Z digest=sha256:4787b07360102e6cce99197304ddbd3c0297483a7b8377aee1e463ae7a7ea84c

Observation fa6ac10c-57da-444c-8232-5b6303355bab · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.115856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.115856Z digest=sha256:48f5d57ca21576b3e671e7e3e8c6c105b9d6f200bddd06ecede87ea499c66421

Observation e61f2bfb-3c1d-4de3-9c5c-a0b9edd41418 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.206331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.206331Z digest=sha256:bfbcd6113a365ac3fb218f13f6f75254181498aab5de47763578671d97947b87

Observation 967f09ed-0582-4c9f-8691-c1c707ee491e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.317950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.317950Z digest=sha256:6b7735838c81b83a822b7b2968fbd967c54a701148667d573523ea78b4e4fa47

Observation ae5e1ace-dbc9-410d-9784-e1158d796821 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.425233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.425233Z digest=sha256:8155509c0ece2cce28f919c295d5b0feb60c856f0b66c6b615ad64119b055d64

Observation ea82e587-f15f-457d-b9f4-c78629abcb86 · outbound

This paper cites Kimi-VL Technical Report.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.543542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.543542Z digest=sha256:690b16c5bdc9a0f96ff1cd6530b90e6c958ac81962a86ac784dab2d3316e7deb

Observation 5214968e-aa2d-443b-8d4e-ed8998ede8dd · outbound

This paper cites R1-vision: Let’s first take a look at the image.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-vision: Let’s first take a look at the image

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.658024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.658024Z digest=sha256:5c3798a90a1cb6ae43dc71f7c28ebf1fe5017a1b300383e61bd8494dbab9c240

Pith citing papers

Observation 9161386b-c4f9-478d-afcf-2bc15d5c4f0c · inbound

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization cites this paper.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T19:50:33.810618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:b486d9ffe24cd9f4080f9eb6cc4cada6a5ac1ab804857bc00459d95610dc3ca1

Observation 5c9dc37b-4062-4914-bf2e-5acfc1754921 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.085860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:92f9899821b27cc6774c6f71784d5c32313e3c3164bc1dc65927d17a33c6db2a

Observation cc3102e1-4c8c-419d-be8b-9b6a5b0169e9 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.532560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:2409a27918d2446622332346e69e83882ac08621efea11db76560f6655137eef

Observation 7fac40b1-5516-4982-9ade-633169de5a16 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.575958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:ce4d5c7ad3a313aa10a6e58bf25329fac354c3315b98cbfa717f48fd2f96430c

Observation df4417a6-4e26-4106-8a73-5ae38a3056af · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.604506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:635f31026e17faa120bcfe3e0b60d1c531712283c45e7c8977ce0abf45b939f8

Observation 42e34c38-0cf8-4028-9238-765ca4e6a02b · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:29.388556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:a7d71205a10706fda32ed34ed8d64305fe259f15df82c903e43aaed39b028d87

Observation 950009ec-c769-48be-929b-79169082bc1a · inbound

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training cites this paper.

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:36.967214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:14:45.547957Z digest=sha256:710ac539478da4c623f70d86cddcf4810d60511aa00e1a167afca67fbe18588a

Observation 91d4c9ac-a3cf-4d38-9d23-bef70906eb8f · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.285898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:2a0cc821a9da223bebe86b2fe01431c5a3cba2960450c6d579f91eda928dcddb