Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 100 of 227 outbound references and 8 inbound Pith citation observations for arXiv:2505.18536.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18536 v1

Coverage vector

measured 100 of 227 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:17.658024Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T14:14:45.547957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:10.284267Z

Reference resolution

100 of 227 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40097270-201b-4671-98a9-96a86ad89f99 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement learning: An introduction, volume 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.011412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.011412Z digest=sha256:c56b7add7c3e5c8a2da2cb80767c0272e45f9098939a439fbf694b4d9212102f

Observation cb105126-c838-4241-9d62-ac96dbab593a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Proximal Policy Optimization Algorithms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.103283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.103283Z digest=sha256:ba8158cbde14abbe533548d43779da01108beb6706cc88610a36906c0dde70b8

Observation 92fa4f2c-49c8-4842-beb4-81fb763b8b74 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Training language models to follow instructions with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.256788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.256788Z digest=sha256:d1d0b649e7a2cd4b83208da6223d1fce9d58a9ec6f3d7e50dcd31f87e6346e18

Observation dc18dca8-c084-486a-bde8-b13d88490781 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.367729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.367729Z digest=sha256:c243d0a2210b26947171e331b87eb04743ec3f01d2cb39b72fb28f9f7b052dc6

Observation b69da4cf-d149-4c1b-96e6-1f6d07b198ca · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.535029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.535029Z digest=sha256:981e2bdcb9ea0898fe9cab1746ee1a031011048e9c1c95f9b2e7bc9b59b3886f

Observation 3ae63e3d-924b-4ddf-9049-9c858c37329e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.639365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.639365Z digest=sha256:cb8e4dbdc8fcc87886a2e4b9c82416f3e0269f0ee934bb9b36981e6bd49dae81

Observation 81b9a938-15ac-413e-89a0-9c18cd2a4228 · outbound

This paper cites Introducing openai o1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o1

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.792427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.792427Z digest=sha256:e7ba0191131ea6298cc7ea4e48729cdc453813e1df78dcbc0b2197d319e5bb2f

Observation 775fcef7-b5f1-4507-a9ea-236e8404c9b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:08.923108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:08.923108Z digest=sha256:18a42c84c2322d085eed7d61e3365b92fb960d8572809a449452fa2ce562e412

Observation e9572b1d-50da-4296-9599-ef43d6891972 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.013331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.013331Z digest=sha256:5d66caae783589e1f27e59bd339cf8c191b020409620d7f6f0bf1520b54bc0a2

Observation 05c33b26-e119-48be-820e-93c2b9a5bca1 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.127322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.127322Z digest=sha256:b40b64cc65f9cf2ce0e4ba2be52fce9ef78d9e163e88c6958416a22f822d1e3a

Observation 1649ca34-a81d-4b4d-9a27-1477797c16f6 · outbound

This paper cites A markovian decision process.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A markovian decision process

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.235317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.235317Z digest=sha256:b3ca05e47ac16c89238b6ea14ff1a9cc903f269269eec2ce3790230d665e84c3

Observation c754e2d2-b518-4e9b-a3cc-22f8c68a7543 · outbound

This paper cites Convergence of q-learning: A simple proof.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Convergence of q-learning: A simple proof

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.350959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.350959Z digest=sha256:8b1e404e2195edf1c4fb2111559c7202d4165ef70a394b7567563fba90439415

Observation b5172b4c-9916-4ecd-89e9-3b917ef718a9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Playing Atari with Deep Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.496847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.496847Z digest=sha256:4b29b583b3bceaec5c40221a0ac3aa26e6b6f0c8bf188cf23b864107fd7ba5a4

Observation d94e0d25-b8ca-41a3-a2af-4be34aa47d87 · outbound

This paper cites Human-level control through deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Human-level control through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.619866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.619866Z digest=sha256:2943610f49dc9e0fb0165bfb301c880b12cf359d8e070187ebcf75b95c02f12f

Observation 245fef20-d355-49fc-8b71-b4eb0183c79e · outbound

This paper cites Deep reinforcement learning with double q-learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Deep reinforcement learning with double q-learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.734580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.734580Z digest=sha256:bde8417c2580bc729e6446a86aa08ed43df22c851d5f08312311ea6b9cc475a4

Observation 2a4e90db-6a84-411d-acba-9da6c3092262 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dueling network architectures for deep reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.823430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.823430Z digest=sha256:41682b64237a1da4a55fccf946377ab79f8041e48c797a8a0a75180fe9dda957

Observation 549c1f6f-022e-4745-9135-a4973e29a50d · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Rainbow: Combining improvements in deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:09.948269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:09.948269Z digest=sha256:0d06b3ffc8e1efb4aa9c9a0a42b1bb499e722faf102fb485dd09b6aa8d5ffbaa

Observation cf7111c2-ba27-468a-8ef9-5abf2e9bb760 · outbound

This paper cites Policy gradi- ent methods for reinforcement learning with function approximation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Policy gradi- ent methods for reinforcement learning with function approximation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.081908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.081908Z digest=sha256:4cddc7445b4e65e8fc1e6621cd3316703b0b092d874ef24d5089e34afe14202e

Observation f9cc4bfa-6b53-4b20-84ba-52b228655f65 · outbound

This paper cites Actor-critic algorithms.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Actor-critic algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.143958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.143958Z digest=sha256:6f89bd4e7fe52e1c474152483fa5372461a9ffbbd8fc8a3ec2697661fcc17c25

Observation e535d775-dcce-4472-9b7f-d40e5c2b0f2e · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Asynchronous methods for deep reinforce- ment learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.255173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.255173Z digest=sha256:67c07505c705715662af7a2c7c1b8f5e022c9b4dca47853116fef6b4763774f9

Observation 97eef202-39fd-45c0-b930-3f4fc3ddc408 · outbound

This paper cites Trust region policy optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Trust region policy optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.401052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.401052Z digest=sha256:f34ea32caa18557125719d32331f6b9e2012ab2df948be30294989006717149f

Observation e44ec016-602a-47a9-9e9b-e07d14cb0a49 · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.524240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.524240Z digest=sha256:e80f5d4ef54b72066b8570e1c6eb458b43f390eba82c7ab717cc894fcf6cd4f6

Observation 18a603a5-fbfa-4f22-9856-bbadb950af4b · outbound

This paper cites Reasoning with large language models, a survey.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reasoning with large language models, a survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.596496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.596496Z digest=sha256:50ddfa0c40ada4c2674226b8cb07d2255d6d82ad19ddac45e1909209ff7aa686

Observation 8130aecb-d762-4666-9385-a6c66663eb02 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.735951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.735951Z digest=sha256:ae2d1379ac5891764243e99002df7836a72e705028e91950559160b631303940

Observation 2507e31d-74d1-4d0c-9646-f009e4797694 · outbound

This paper cites A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.828360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.828360Z digest=sha256:f5819394b2c4e7bc9f21ab183f87237a87893262d72afff3e2e4277d4c242643

Observation c7d0d5d5-e971-4419-b65e-6fb3f6c96c97 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:10.991423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:10.991423Z digest=sha256:4de36ab3bf0f772c3b964d0c1824cc927f11b4893f6eb358a2c9fd59e6b933cf

Observation 4be25c8b-00c7-4257-a608-0a4dfca4605f · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.116827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.116827Z digest=sha256:6416cc5000c8268ffdb6a07e4a2d2ea0d5c7557a8698f40506e961e7907eee6b

Observation 41b48aa1-d17b-4230-8db2-f10c5ecc337c · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.242194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.242194Z digest=sha256:8e0d4f5a61833dc71c994afda9ab4efe94979efa396dd12f6bbca85857a72a12

Observation f2629d68-8aff-47db-b431-dec3ae56af11 · outbound

This paper cites Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.393683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.393683Z digest=sha256:8ac50aa17b4b0947b88f465d74ddeb282ff23e0df21743c9f68fcc31e898f6e0

Observation 7babe830-f6d3-4d4c-853a-38d3ad66ec9d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.547423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.547423Z digest=sha256:f038ca025f3f01bfd5261a1467458112d3fe55427f2dcdd3dbb5cfc6b011156b

Observation 7b47eff1-c366-4880-b0c7-63af299acbcb · outbound

This paper cites Perception, reason, think, and plan: A survey on large multimodal reasoning models, 2025.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception, reason, think, and plan: A survey on large multimodal reasoning models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.658758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.658758Z digest=sha256:280e1294f421cf4023d2568b9c2b2ed3893a0501a527ebe80be6af9962e3be93

Observation 1580e653-fd74-486c-9a03-a89e1b2efe30 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.745517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.745517Z digest=sha256:31250feec3b94b6e518ba32c5d811ff40cba89b0dc1452aacbbe221dba88d9bb

Observation 1c63fa0b-ff1e-47dc-ba9f-59a724e5c016 · outbound

This paper cites Introducing openai o3 and o4-mini.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Introducing openai o3 and o4-mini

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.857829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.857829Z digest=sha256:9ec16aa50db6c86c740f74f91b1009bcd5a291a5367466824960425a3a308f7a

Observation 16623895-9b4e-4516-b4a9-2942d29a3356 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Grok 3 beta — the age of reasoning agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:11.968928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:11.968928Z digest=sha256:9f0cb3daa88529ccbd927214a49991653144fb944fe1486c3944592604cd4587

Observation 86bcda32-3c84-4331-a4b3-9b6ad94135f3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.068390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.068390Z digest=sha256:8f5f63b037efdebd3aa4191d4cb8c75309857dc780ce7d031689c2e740eb3007

Observation 5a962b26-e270-435c-9054-635f65b65d34 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.204895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.204895Z digest=sha256:660da3f7f9cb0254db15c7aeaf959e2f72afe6b21b0e47ed672b2ceccdd362d5

Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.361455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.361455Z digest=sha256:19fe1f17ff787c301668f027f3daba15f9bc9d61148720c45b35afdc75ff089c

Observation cec5c169-3d82-4565-a809-74ad2bb3a543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.476155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.476155Z digest=sha256:88b8b46383a470e59f21bfd870dd1222419d0d56910d398eca3ea4fe0ba3d2e8

Observation 3deeddac-e41a-4e02-8e2a-6e3b1a0ad8f3 · outbound

This paper cites Approximating kl divergence.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Approximating kl divergence

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.595428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.595428Z digest=sha256:a420453186e8da1715d2085e518714ab7847c3e708297f0af27c9342da767c19

Observation d9a22cf9-07fd-4372-a117-a775ac9d9bc1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.705003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.705003Z digest=sha256:e63a8646cc5fd627087cc2e7e02adbea69327fc4093cb55c4f6a3adbb98477b5

Observation 19d230f4-4c81-402d-850b-fcd968b02023 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.807492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.807492Z digest=sha256:01e3c217c4a33440d9e30805cd7c7d9d8b2d850ed2c983e65bd8a1da87aec305

Observation 58bcbb80-eed5-449a-9c08-cdb14c4ad7e0 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.891931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.891931Z digest=sha256:f79dc724868e1fd1b2df1d6e4d4ad0c2ca3401b7ef381b103516a1b118f277d8

Observation fa8e1e95-c984-4d5d-9d37-e36825a8a940 · outbound

This paper cites Tinyzero.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Tinyzero

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.022145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.022145Z digest=sha256:14f5bdea4a2a302bb7447104795586ff88fb04a369ad31adcab9337887cb71cc

Observation 061d8941-df23-4968-95e9-82c3a7d639e5 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.154734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.154734Z digest=sha256:c88798c758e853f67a7645819dacb5c253cfea2f48be65e535393aba563d9329

Observation 91ff927c-8585-4733-abf8-33827210c667 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.239675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.239675Z digest=sha256:68f1e98e33eac5458cb28bb7f42dd9cd5df65ec96980c31361d96b880649d008

Observation 051c8f17-2752-4dd2-86b8-fe78be8377d2 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.310672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.310672Z digest=sha256:cd710b315e8fdf248d932a6fa53a8acb5a29daccf806f98a238de7da4e3dbe31

Observation 65bc45ee-a926-45a0-a762-ba44de584247 · outbound

This paper cites Audio-reasoner: Improving reasoning capability in large audio language models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Audio-reasoner: Improving reasoning capability in large audio language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.413579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.413579Z digest=sha256:9f2fd07f25f6da605ceaa4ee704cac092953cccd12fb77566df7bdaa7201a930

Observation 7b9ad463-7fbf-40ee-8a5f-8390c59237f0 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.461419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.461419Z digest=sha256:e7e519fc3106227a6f1f7e534551fe718596bd54ed47e5fc08d4057ad8748c15

Observation 4da64ec5-3e68-4166-89e9-1a7e253e5cce · outbound

This paper cites SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.506592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.506592Z digest=sha256:17a2927ac4236c3e340c84d868cb7bc0fdd6548e3390b5286d3451f9acf2f877

Observation 0e33bb8d-c9fa-47f6-aa5d-e271e2591963 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.565978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.565978Z digest=sha256:ec8d5b22714a29b4acd04aed132bc4facc9a9052ffdca91d73ca9f7475a8a211

Observation bb8cd184-7364-4401-a920-23f1239dae01 · outbound

This paper cites EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.645079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.645079Z digest=sha256:4300780f61c09107f9c4a25741c5d865fa67bfbb694aaed44b434394a204e661

Observation 5c1d1cfb-edff-4ec0-b4f5-2733b3adcbfc · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.700122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.700122Z digest=sha256:3f64b46fb9e382d845c8f6c370c858ef3ea3160d727c542e057ea42da3461c7a

Observation df3377d1-a140-4ef4-9cf3-330f7444868f · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.747182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.747182Z digest=sha256:160cd4addb109af9ae62372c2446913f66e273e2b86189ebf2747a02216d5cf6

Observation 536638f2-8005-48a4-aa21-835ba8a79f81 · outbound

This paper cites InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.797112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.797112Z digest=sha256:c98ad6f9822d26562bcdce0fc2ff43fabeb162178469b40b89f8005b95321132

Observation 3f8fc103-ae94-46b0-a2ee-6afb7aa41df0 · outbound

This paper cites Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.851407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.851407Z digest=sha256:be281db1c7733585bd19057bc766430a4c2b718260ff8fa1ebe06a885a72a15b

Observation 5bba1d4a-7bf3-4a4a-86c7-0cf2c789b572 · outbound

This paper cites Vagen: Training vlm agents with multi-turn reinforcement learning, 2025.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vagen: Training vlm agents with multi-turn reinforcement learning, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.898943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.898943Z digest=sha256:2b8179ac654b92b173c5025010d43463ec96cfe858469a4a4e6184f2e9fb3c48

Observation f9b6a6df-454f-473b-a4f1-bc38cbc032f4 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.950384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.950384Z digest=sha256:03c43ce635368f6c3a94596402e78d872bbcf34cc27916400a28364e4f675ba3

Observation e804e79d-632a-4614-aa32-b12fa31e5490 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Measuring multimodal mathematical reasoning with math-vision dataset

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.994703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.994703Z digest=sha256:a6d66155d5641ca324b4bbd0340fe60093fcbc32ac0a3d76e065a61b9054aa58

Observation 8ec617e7-3684-4a89-8b41-a39d5fd66cae · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.050530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.050530Z digest=sha256:328278ed2013fc99772960cee23ef1ee93dabcd62fe202f333ee9c50bc5ffd7b

Observation 081cfcc6-85eb-496c-9b7f-2c9ed819fd1c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.095636Z digest=sha256:6dd867393e0617271a46f719645d97bfb15946728f8bcc17a632f4bfe97b5e91

Observation 7e0cb03e-51e6-43a9-81be-841156b56272 · outbound

This paper cites Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.148793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.148793Z digest=sha256:2cc7e2240c449e8f800ce7f93d5e161ac6827da6501bbb6075aadbda181bf8e3

Observation cc5df7d0-5c0f-4aeb-b61c-29c32bf1862e · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.198563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.198563Z digest=sha256:03bf8e3303307526f2575648b7352821e228b68947aa0bc48fb56b30e7175727

Observation a04b69c3-4337-40da-b1e5-a0d8569631fb · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.248813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.248813Z digest=sha256:462106490d11412e1bfc123fbe13b27b58aa68941a1e4b4eb463f65d241c4c0c

Observation 676d8967-aa9c-42ae-b9b5-db8c76bef781 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.353463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.353463Z digest=sha256:6566de41619f6c8a671ccbab42ccb087d3976760a640a7e031767bb21e52cba7

Observation b2a65190-299f-4f0e-b539-ec674f8f4100 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.410946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.410946Z digest=sha256:a3419c17d831aa8a3d6e9c14665268eec782f300756cdfe93de06116533cf9a0

Observation e892892b-c10d-45a6-8065-9432d3feab2f · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.468895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.468895Z digest=sha256:1ee9f2d0303f39d7d595e0478fd00d3ef94ec1431c1de471fb9993108c87b051

Observation 72cdbf10-5872-413f-8e80-172e4bc18ba5 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.524790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.524790Z digest=sha256:32672d91c0746edc783dc3dbebbb0079fa1ef94f07b6af5e0829b5a5b07ed15c

Observation 9ecc574c-ada4-42b0-9d5d-6131135f8e9a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.568519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.568519Z digest=sha256:ce356792e31ed3d189e2157f26dbad8578f9115e35721591be595b6b1a857393

Observation 1e676ee6-46e1-4876-99b2-0bbc10440dfa · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.613909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.613909Z digest=sha256:8039cdf9a48e6f6b1aa7263941b60cc71aec555d308bc94e25dc22b79373782b

Observation 9fa25072-aa78-4ace-88dd-3202fa3459b9 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.672507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.672507Z digest=sha256:82bd947012957ecef3a1b1ba4cd628f8bd1b6b2a6922aae3e69090e4bf2b55de

Observation 79d184a4-a0f3-43e3-a7be-ea742246e1a9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.730553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.730553Z digest=sha256:35120382116d54260bb9a9d510d75e47f23d22e41e3027b8f8c5dde7c86a8c9d

Observation 5f5aa028-7c61-4799-ba77-a4ed484c335f · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.783221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.783221Z digest=sha256:3a661c394e901af157863fc5c9cb74f9073ecb2b041911caae3bd7f32cc08044

Observation afac0b4d-6a62-47bd-83f1-497cc0567c8d · outbound

This paper cites Mmr1: Ad- vancing the frontiers of multimodal reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Mmr1: Ad- vancing the frontiers of multimodal reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.842854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.842854Z digest=sha256:5ebb712bb0e99f1d74d5da553352b290f8ec0a42584db4df2be033d0c43d9c32

Observation 1f12bbf5-1ba4-436f-8aad-46720d7a2ad2 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.922706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.922706Z digest=sha256:c4deaf67c10f6e024a58f969ef2eb4defd75b6c2ae442e2034cf5443f008a614

Observation be1b0c93-c556-486c-8a10-b029c698a2e9 · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.986136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.986136Z digest=sha256:f65b13fb5c50bdb04f0ea413c92762139054d5f7a1280b111d210bc21d27695c

Observation a457b09b-f3bc-4a6b-acf1-bffc6f8216c3 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.058067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.058067Z digest=sha256:edd1185f8cabd2a2df22a7edace7659af78e10092e1d5e3dafa9dadde11d31f8

Observation e6ce94ea-7263-4968-aee8-9c4a75542481 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.106642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.106642Z digest=sha256:b19c0823bb38c1f8214812c92fc5cc2c41e41f8bde1200e0ff409469f860af60

Observation fab00f3a-ea59-4597-8b1d-393e34eee0cd · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.147711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.147711Z digest=sha256:3d1df37a1f5c019fbb842c783999aca68348e5f3d3e539d985d426ff70cffbbd

Observation 5b345cae-a46b-427a-bb86-0ff66b6bf849 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.223794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.223794Z digest=sha256:1bf7af9531940f07637ff75d14a33773e216e8fe2170189babb822a0d59e4b29

Observation 00eac3aa-c4fa-40f9-ac6b-2a9297bbe228 · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.321786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.321786Z digest=sha256:d5e184f0bbbefdae0d76987de2c35a1c6f3e6542d2b3dd6e65708ebb526201b8

Observation 4a7a18cf-e7de-4faa-8a78-5e79f51dd4f3 · outbound

This paper cites Fast-slow thinking for large vision-language model reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Fast-slow thinking for large vision-language model reasoning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.448124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.448124Z digest=sha256:6dab0c51ec76bd7c5354d8125b70e97ac47aeef6c926831b457030d32365ffd0

Observation a1036505-1980-4395-87a2-f3dfba3bb0fd · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.565534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.565534Z digest=sha256:06c872b2f3bfc4fafa48d99d248ce1b8e85911453d62a38887b75e69bd1a6e15

Observation 56723f23-61f3-4feb-a363-1b8e79caf6f7 · outbound

This paper cites Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.674647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.674647Z digest=sha256:076cabd08864fec1f93dbefc3092d82da2d2fccd005af8b7dd74fb8e38cbd7d7

Observation 570fc3a6-5555-469b-82cf-12e948096f05 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.750218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.750218Z digest=sha256:163d3f17d8caa9ddb2b48a4bd5af25b762170fb32cae79714585225c7a3a5d97

Observation 21a87c45-00b5-4cc3-8c36-faf3d4d6f4e7 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.837024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.837024Z digest=sha256:2a297ff3f9efab00d11851dce6a929170e850b5714966a9c46512eac54ac570c

Observation 0eff6400-1167-46f0-8655-a051a7ff64cf · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:15.937329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:15.937329Z digest=sha256:daed08e50cb4b8322bd3b22427e802b1e6b65b767def8728f413d9b8df96440b

Observation 496011d8-4e0b-4c11-b078-090c603f2b73 · outbound

This paper cites Q-Insight: Understanding Image Quality via Visual Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.030204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.030204Z digest=sha256:b126b79e23ab6a985f1c15d8398478f55b4d478e44d1994069ef2f97357dd519

Observation 54ee01f6-ba4b-4c0a-bea4-e51e7d7f9665 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.116604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.116604Z digest=sha256:dc422ad63fd1feedace32087a791552bc15649bbb19d2a8e66094a427d4b0c75

Observation dcfc2b59-51ec-41ed-b2ca-a510e7d13a85 · outbound

This paper cites Compile Scene Graphs with Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Compile Scene Graphs with Reinforcement Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.244853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.244853Z digest=sha256:41cff3d6273db3bdf1f49ff38f35f1ef22cc96b9f2f10835eff67ac9f6725ae2

Observation 6fe23011-cb10-4f82-b0f2-5d19a96a19ce · outbound

This paper cites Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.408178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.408178Z digest=sha256:0ae0d9d2b55cb1b24b110afcff1bb9e59a68a96fa417a00488a13c5fdf313144

Observation 1eee2c39-18e0-4c37-8fa5-68c866112ea6 · outbound

This paper cites R1-track: Direct application of mllms to visual object tracking via reinforcement learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-track: Direct application of mllms to visual object tracking via reinforcement learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.533191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.533191Z digest=sha256:be31d9bb9d68df7a7c37d049bd2a711f8a7a85f6dc08e2f49a27be1a9cdd724a

Observation 5cf62a27-321c-471b-b20c-074f2896c833 · outbound

This paper cites SeekWorld: Geolocation is a natural RL task for o3-like visual clue-tracking.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SeekWorld: Geolocation is a natural RL task for o3-like visual clue-tracking

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.665298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.665298Z digest=sha256:dc2ed9b09d819089c291fcabb0963e97d865817beba8a845475dabe05570cde2

Observation 8533d11d-b34c-4e6c-9ea3-fe6cdb35c3f8 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.855930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.855930Z digest=sha256:96ac0d691994b8beb36ff682517c3a818d3253044471114a2eddb7d0bf414adf

Observation b86b18bf-b91f-452c-bf36-547c9fa8a0c5 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.974601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.974601Z digest=sha256:6e62f0e9b50c4d5778e165e98e7265c2d0d30ab70639019d8dde138892d5c50b

Observation fa6ac10c-57da-444c-8232-5b6303355bab · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.115856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.115856Z digest=sha256:4b9a57136f467e3fa2a6e83eac9291491a9cffc458a2b3e5c5046250014c115f

Observation e61f2bfb-3c1d-4de3-9c5c-a0b9edd41418 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.206331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.206331Z digest=sha256:bb0ac521a4e8fff70ad9549c1e49a48e7df99099d39938b3a6ff0e1d266bb2ee

Observation 967f09ed-0582-4c9f-8691-c1c707ee491e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.317950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.317950Z digest=sha256:e701add6ee78908e1f3b85d0ed75b04ca1bc78186c24cc8e1110cf0cdb98b91e

Observation ae5e1ace-dbc9-410d-9784-e1158d796821 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.425233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.425233Z digest=sha256:8bcb49e850e88462359f7e9ddb74c5541260d20f9c3584abf0a26f7aa414fa43

Observation ea82e587-f15f-457d-b9f4-c78629abcb86 · outbound

This paper cites Kimi-VL Technical Report.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Kimi-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.543542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.543542Z digest=sha256:7d0a8da393d1c6b97d088eec71752635b39be0b9fb6e88f473ee3aade7597810

Observation 5214968e-aa2d-443b-8d4e-ed8998ede8dd · outbound

This paper cites R1-vision: Let’s first take a look at the image.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-vision: Let’s first take a look at the image

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:17.658024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:17.658024Z digest=sha256:dc21e00fa76233a4c0204b77a4ce277775861468a8db242caba0db9da688b557

Pith citing papers

Observation 9161386b-c4f9-478d-afcf-2bc15d5c4f0c · inbound

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization cites this paper.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T19:50:33.810618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:3b00cae3176cbe3da619e8bcca1da594344f7eaf5c3c086a76fd918df98826f4

Observation 5c9dc37b-4062-4914-bf2e-5acfc1754921 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.085860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:fdcdbef29495d7d9531bbc2b11e0638baf8c0b0b0777a3882ccdc0cd92f5ee2f

Observation cc3102e1-4c8c-419d-be8b-9b6a5b0169e9 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.532560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:188e4282c3a1a263637fe953b8e676e08b0e8705a474fe2a689c99e4ade1eaa5

Observation 7fac40b1-5516-4982-9ade-633169de5a16 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.575958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:2f3b95032f7d229c1a40c5ebe2e0177c06d4ad155fdec7cc43d00d97084d9580

Observation df4417a6-4e26-4106-8a73-5ae38a3056af · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.604506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:f3068b801224120136e47f0e0e1798d52e41df7985c7aa3d0e588854847de561

Observation 42e34c38-0cf8-4028-9238-765ca4e6a02b · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:29.388556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:15ee732d0543b63028a00552e15015050f0199eb2bc86effce7da34704515e8d

Observation 950009ec-c769-48be-929b-79169082bc1a · inbound

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training cites this paper.

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:36.967214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T14:14:45.547957Z digest=sha256:1d0ee1cd9f5da312897325c5d86fa9bc552ea0b5deb3d9c0548aeb542a8952ec

Observation 91d4c9ac-a3cf-4d38-9d23-bef70906eb8f · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.285898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:8c0d45020caabeb32f060bf84e6dddd749401a669d1e6e564e0e229b7f183040