Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

As of 9 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 10 inbound Pith citation observations for arXiv:2505.22334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22334 v2

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:01.038049Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:47:07.116263Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.732932Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf2f7f32-6e97-49ea-aaa4-86384daa567f · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.537010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.537010Z digest=sha256:2fc5dcfa3f988dbda5781abe1f3f265d383f6469844c477eab6ae6b31c6cd366

Observation 7282e936-d0fc-4c87-98b4-c92ea2a75876 · outbound

This paper cites Qwen2.5-VL Technical Report.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.627521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.627521Z digest=sha256:4f96f1894ac3d3c31265b0c31840c40102ba9ec150671b125406c464e8119201

Observation 7aa782b2-efae-4fe7-a425-fab274a35187 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.685659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.685659Z digest=sha256:c8675d46e36561e6f13c28f21231b64381bb2740e300818cc5a8cb62283beca0

Observation 8ea13821-0f53-4f4b-8ade-72e0f46458a1 · outbound

This paper cites Sft or rl? an early investigation into training r1-like reasoning large vision-language models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Sft or rl? an early investigation into training r1-like reasoning large vision-language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.793186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.793186Z digest=sha256:28035f2043d6b2de68d14d6725d5dcc53971f7f7a5d1617b02ad0bb9e3bfae36

Observation c9b8d8f3-4bab-4c2e-8cf4-ba4ff0f32a3a · outbound

This paper cites Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.892662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.892662Z digest=sha256:923039967ae3d4532425c71609ed1c083455325b9262b55204a631b8e9dd7f3f

Observation a3ba1c2c-06f5-4181-8bac-377a0ebf08fd · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:52.971155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:52.971155Z digest=sha256:da26912346efaeccd08e1013a23e28849d3c4c7dd07a5e2741009ea0d2cfb443

Observation ddd47ed3-9fdb-42bd-b6ce-cd2a04bbea3b · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Sharegpt4v: Improving large multi-modal models with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.054692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.054692Z digest=sha256:c2601245d6743ffd892fde256318f2b1049971b757861d017215342999e4fec5

Observation d6e6e4a9-3d92-4866-95c0-a526d30c03aa · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.142788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.142788Z digest=sha256:97abb85be47bc51db821c6647a887b15e5501a4fcf2c5f28a2e09292f95ff2e8

Observation 0c35fc60-7847-4323-a846-5c23a3736f1b · outbound

This paper cites Vision-Language Models Can Self-Improve Reasoning via Reflection.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Vision-Language Models Can Self-Improve Reasoning via Reflection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.239198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.239198Z digest=sha256:2ac89b04f784baa161b5dd18f57ea99b8c92bc4158ef06d7ffc3112fb5ec7e28

Observation df2bc233-7654-44c7-8989-c034770064d6 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.300884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.300884Z digest=sha256:0888d623b3c0ede7c39b903ebf841a2923932f5ac0af9034073ff38325b18077

Observation 33e31527-5aaa-426e-b768-4588cb091ccd · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.376458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.376458Z digest=sha256:79e8b97d9f97a9dcdf589d4b04d791d561301cb7275a3b0d35beb368cf118449

Observation 3e83f93d-a0b6-4051-9d6d-f315e394bb53 · outbound

This paper cites Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.471122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.471122Z digest=sha256:d5fcd54f1051858d7016fb4df17335b0e5758c813502b279382f20dcece7d4ad

Observation 4d36db9d-6f0f-4569-8293-21758f36f9e8 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.564403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.564403Z digest=sha256:f200a4eca9ba08427542286b43dc726a53a8145b21301b3d54f1651422e57836

Observation 42f47a43-0c55-4d9a-87f5-69e2caac04da · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.642938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.642938Z digest=sha256:d27c5c8040cb51e3c267e80eb6d121277a7911a1f6b2f7e06d1d9f61229d2956

Observation f099e45b-584b-4429-90f8-7bfef4ee4b09 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.708467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.708467Z digest=sha256:5eb40cc45acce5c9f1e197d6ea53961df1b4d34acf0210aec33a61ef6c01acbc

Observation 1e8e55a2-7f5b-4e2d-82c6-7bbe078069f3 · outbound

This paper cites Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.764326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.764326Z digest=sha256:843d684a407aef9bb940f850dc5160bc51ded969483dcf7bfda06aa456b482b5

Observation a7706e48-a2e3-4ff6-bf07-642cdc4169d1 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.817955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.817955Z digest=sha256:c0c67902ecd3a46f745020e239ba8b993f2a6251f1943d8e470a680dcc90e225

Observation ffe533d3-e574-461b-a2d4-fdd83d57c5ac · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.897667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.897667Z digest=sha256:5271d1d2009b3c42fdf370c8b55cff4ba61a545790600842195067c2f430a42d

Observation f2d1fc8d-40b8-4614-9802-4c1dbbd23734 · outbound

This paper cites GPT-4o System Card.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.965180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.965180Z digest=sha256:934b409090eb27e4d805a9ee1885fbfe5ad0037bd775c30093446621ca1fec71

Observation 35201955-b0e0-4f7c-8e2f-65c9598af8de · outbound

This paper cites OpenAI o1 System Card.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.061385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.061385Z digest=sha256:7448ba2aa51a6b61186d8e9a70ec8527131fe7523f5d99fb67ccbb16e87bef03

Observation 3d7fd054-2c18-4304-8e71-800a895a9c70 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.135029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.135029Z digest=sha256:4fa1f5c0e481cbb72f06431d3923c3c0c425dedc0783aee4489d5f5876e412a5

Observation 2a89bc6b-2403-405c-b6d1-3e3b02064ecd · outbound

This paper cites A diagram is worth a dozen images.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start A diagram is worth a dozen images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.181520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.181520Z digest=sha256:812e4d4e8e89bb51ea8b68f29cb9ded44cc04d046955c82a5fa66cae00b27ee3

Observation b394d188-35fa-4666-a776-f1b32039608b · outbound

This paper cites Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.261687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.261687Z digest=sha256:f9febaee3776f2c5ebcba2843815ef2a13dd857477e15842ca796cf5e336d739

Observation 0d004864-0200-4949-85dc-6c912af1fa89 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.353238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.353238Z digest=sha256:a449ce6e4d42c1d028c69bbd304849d88a2cc85e2c5b0b21ac0857b4abbc023c

Observation 037b8c57-1954-4814-a83f-344cccb655ee · outbound

This paper cites LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.424523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.424523Z digest=sha256:0c86ad6546184af6a0ba1ecf28e8a1cb840f53edd59459a48127183928199afa

Observation 9a73e6a4-0672-489b-bf73-a3283f55e690 · outbound

This paper cites Small models struggle to learn from strong reasoners.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Small models struggle to learn from strong reasoners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.501039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.501039Z digest=sha256:dbfea73b8f1395b0e9b30a6fc98478150de85aacf65f1b1b2b5a9e243fce87d9

Observation b417db3f-86af-4854-94a3-596d2ebdd857 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.565717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.565717Z digest=sha256:bf00a6a9e9854518c60d12d0483b307f8183dfca679df2ec3472df078c202e3a

Observation 17c6a821-cd57-45ff-bd17-2baff1321b62 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.641119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.641119Z digest=sha256:059a84d8df2dcf0c75d0cf9e099c583a659f3bc096ce944ac0393f698bb12847

Observation 97844a3d-0ffb-46b8-be1b-0f8825a089bd · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.713753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.713753Z digest=sha256:b14718fc718d3af1299c9966282c06977431f18ccbce63b1f486164ad4539ba8

Observation a1f7142a-71bf-4066-af12-08221de474ae · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.862278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.862278Z digest=sha256:e20ce2f8ff27974818d670fef94ed62d9dfcc4cdb36c2b15ebc30d1bdc8a7dc0

Observation c06c8f85-6815-4e2f-a344-637fcd7bc9a1 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.957538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.957538Z digest=sha256:45b82275443df1ecdebba3222c4b31ccac2dcaf7742388d718cb1f8427aca211

Observation c734c683-f7fa-4713-8fb2-9c415a15fa6c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.080042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.080042Z digest=sha256:2ef195311142774a89c4bb84bb3a8714efda295f7b3406122e811ff16b097f8e

Observation 03176429-43cd-44e5-8e28-4dc958ab35d6 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.233108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.233108Z digest=sha256:07c46b164e7e2c9ea0c5537d619e8f8df48e3e80a4f468c32b0713b31c630c31

Observation c46935cb-7cb9-4e3f-b2a7-4f27b43ffc7b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.396310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.396310Z digest=sha256:8b88678473469fcadb480b26919f2ce1984555e1e4b20663844500d6e6d7a045

Observation d30813e0-15f8-424e-810c-c210ec23d1a0 · outbound

This paper cites Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning, 2025.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.547632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.547632Z digest=sha256:3f4cdad9418db01a63a9afb3f9f1bee29f082887d25fa300ff531eb2ad5dfe58

Observation 31def1ef-43ec-4b2a-aa73-17b3d493d011 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.652216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.652216Z digest=sha256:2568b7956b0cd2af245620901182b920103ee17a8b2eb514ac29012fca3a682c

Observation 90ef9e79-5272-441e-8a41-c03e13bc9072 · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.775370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.775370Z digest=sha256:1a1d076ebc7ba59afeb5250272a66238274f7c450c5482d0cf0a4f8f8d54f91e

Observation 186b519e-bec1-4ba0-a171-e04fd64c632a · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.912190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:55.912190Z digest=sha256:a831ad29ea21d2cd30f4ed6c7fd80dc413cc0acfd33b34679801b5e1e3809db1

Observation 88ea283a-40da-4d23-9579-e0f1ed81e090 · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Solving geometry problems: Combining text and diagram interpretation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.022698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.022698Z digest=sha256:1333246986ce789dc72a89d33e5dbbeea6eea6088c78d382d6be532f985a7907

Observation fab60022-41ec-4e5d-863b-f1e9c041f5f7 · outbound

This paper cites Rethinking Reflection in Pre-Training.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Rethinking Reflection in Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.143869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.143869Z digest=sha256:e3e9be764a1d8e45328eaea2c527d46f1b64c5fe028efacdf1ab2c58b326783c

Observation a11a1ac3-a697-4387-97e5-d8d96be010c1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.300104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.300104Z digest=sha256:d39932e5b1572df3410a515347aef601efc3b25d7a25476448893ef67d37480f

Observation 343899cc-4aa0-4490-8d04-4ebe9f2a9052 · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Vlm-r1: A stable and generalizable r1-style large vision-language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.449948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:56.391949Z digest=sha256:c3f024966e8378504fab5de78cd5532021991cfe87e44eec381b44cfec57574b

Observation 06a8e1bc-ed18-460e-ac83-75c63b3af582 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.479241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.479241Z digest=sha256:c7c7b3df9a90260dc1f92b96da6327583f1a1f9fe0b52b7d2403cd2b5e36a024

Observation a10defa0-6cfe-4974-842b-2b1dadf18f59 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Qvq: To see the world with wisdom, December 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.588932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.588932Z digest=sha256:cb2e8bb97afdec45083cd21d2c3f6d2808c0e5ea2f5e348f2bcb3edfe412dc6f

Observation 901117d6-9bc3-4d6e-a114-6f39238faa11 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.178034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:56.709212Z digest=sha256:cc858d81953b373e0fd34507b2fd7e14665c2f61f98b703dcd07ef7f580fa4c1

Observation 5af3b130-2555-4395-8ab3-95379f7f4141 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.792142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.792142Z digest=sha256:4352ca482d66680cec9922bc25c12a50407843f0e49bf0743b4749704786b9c2

Observation a79d6d28-d27a-4143-bca6-a762b35c1019 · outbound

This paper cites Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.898278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:56.884248Z digest=sha256:cd220eb0dda7fbf00377de5e7666c2d99dd6985e868bae60842ac5231f73aa21

Observation 9f4544bb-202d-4cd7-ac7b-b2e7a008a2d4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.989615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:56.989615Z digest=sha256:515fd86040d8fc68762e31d507b0c9de3ccbed6864c752e6d22cfb880a4e1994

Observation c19e7c57-99e1-493d-8590-78bd419dba88 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Measuring multimodal mathematical reasoning with math-vision dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.644661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:57.115706Z digest=sha256:4f123c3f036ee88634026e09c7ee9dea8808872fb188248624dde5b291b5c2c3

Observation 2d3ccf7f-34ea-4d12-9641-731228058c6f · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.485039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:57.257567Z digest=sha256:06691361c3c1dac0b3fc9ecc204829b838547c3d93dc57a6df25fed0646c6fba

Observation 39f7e458-7dd4-4b3b-9084-63c8f4085552 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.363121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:57.363121Z digest=sha256:83bbd7e6e87a4db1cc418c202c1ea8c4ac0ac4f8913a99d965bfbc0e50a892ee

Observation 28084151-8b5f-460c-a165-44b5b6085461 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Chain-of-thought prompting elicits reasoning in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.504024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:57.504024Z digest=sha256:7629635420ee507c9c5e70ccd286fdabfb0fff13f0cf12776780a43f0a5af5ad

Observation f80afabf-0589-4ed3-a5eb-94ac1091de39 · outbound

This paper cites Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.631505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:57.631505Z digest=sha256:94a3d737764b92fb3335e0e6520908f8500d58e15a9aea66cd442244f4deebff

Observation 87d06979-765a-4f4f-b19d-613fc9a935fb · outbound

This paper cites Thinkpatterns-21k: A systematic study on the impact of thinking patterns in llms.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Thinkpatterns-21k: A systematic study on the impact of thinking patterns in llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.789354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:57.789354Z digest=sha256:cda099f428e076c39388bd7fd7a6242cff790cfba79f925a14e8073cf5f7fc56

Observation 170cf958-85f6-4120-a855-822dc9e166d9 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.865700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:57.865700Z digest=sha256:c439cc2b3103b901aa682e59b17323a6580e39cd8fc4baacae76c826a1d5cf2c

Observation 7f2d9f59-0262-4b0c-97d8-8aabdb75a8c0 · outbound

This paper cites Tbac-vlr1-3b-preview, 2025.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Tbac-vlr1-3b-preview, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.313434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:57.964220Z digest=sha256:420a444664a1de23027b6f065f2e3be24f86d4fe3fe7618ddd618ff31f718f15

Observation 7b522abf-9414-480c-adb7-a27c2fc1873c · outbound

This paper cites Qwen2.5 Technical Report.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Qwen2.5 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.041766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.041766Z digest=sha256:2ace6f045bb44a8d14cf24f110c5c028505d1e232de758baf11f8fc228298381

Observation 31ac8f67-bd2b-4af2-967a-3513938cd04a · outbound

This paper cites MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.101610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.101610Z digest=sha256:5690a113b19be6d8b79610bb1fbec6f024f024076e3afeaf5ece14f5a146584d

Observation c410146c-2b0d-4cf1-9aec-ba6358eaa55b · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.218444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.218444Z digest=sha256:7c74272e28abd28e87623f79cf31afb273e53b37c8301e0615b88a0ffba6808d

Observation b5d7a13f-412f-4c12-ac67-fb5593a4cba9 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.041294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:58.305710Z digest=sha256:4d0deb43b628bf9d643f5dad7a83dd05d2d6f61d367d68f70b45bafb9195a5a7

Observation f0892b82-1770-4b4f-ac1f-d978e917377f · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.382458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.382458Z digest=sha256:a826d3d629a0f03132cfdfbc4714d7ed07d9670e00fa2bcf6b6e3476cafe6c51

Observation d68c3269-b1d9-48f9-8ace-48cd4b1d7387 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.487945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.487945Z digest=sha256:cb9ef5e3b81a0ee081bfd81a7f242a908be5995c117aa3de35ebe1fb0d487baf

Observation c040f369-7c8e-4b6f-866c-e0a3ce6ec44c · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.601919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.601919Z digest=sha256:c4ffe445cd2c8679528cc1be53c7b99d68f28f5ef07e44f61305888d751841fa

Observation 25b11409-beeb-43db-9ca3-f480f5916501 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.712192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.712192Z digest=sha256:59c9906eb1c4144fece32cbc0df7626eb464b10415883e382a808082b3674c94

Observation b4a26876-c9af-4471-865b-b57c0d141a85 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.778569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.778569Z digest=sha256:01aac29f549c112e2ca58a69aa60b260e3eac8d6978f71ed024149aabc153493

Observation 9bbde23e-4ac3-482c-8e41-cabfad96a542 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.908022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:58.908022Z digest=sha256:500e29526432ff1f747d173365b66157f6fe161b3e23708a8292b66c7b1ff119

Observation 706c7d72-c48c-4673-9c71-8cb089abe8f8 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Improve Vision Language Model Chain-of-thought Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.008711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:59.008711Z digest=sha256:ca2c9700696b601061966499082ab02fea6631d1f55f567cbd6fb41d1f404712

Observation 637323d1-858a-4d0a-b7c6-8f7a7fdfc1f7 · outbound

This paper cites Question-guided knowledge graph re-scoring and injection for knowledge graph question answering.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Question-guided knowledge graph re-scoring and injection for knowledge graph question answering

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.086822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:59.086822Z digest=sha256:50858cd79e7f05a826ece3c9e0971569cf6b16f8d218ad15abffb25c1e83e634

Observation 5bb1439e-498d-4173-a198-15e309bf668b · outbound

This paper cites SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.184860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:59.184860Z digest=sha256:d7071732fadb3ea819cd23098954a0f1f5d976b246abff52b0448196f088e035

Observation 482f2187-dfae-43a2-af35-2d15e290995e · outbound

This paper cites aha moment.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start aha moment

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.873071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.280504Z digest=sha256:c1dc5f2b46fdc5646edfaff99190231ce8a2d80ff310c1739ad3d1d87ade83bf

Observation ba021b73-5370-4339-b57a-0d1596a8bdde · outbound

This paper cites Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.356894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:59.356894Z digest=sha256:61652424b5c893cab639a76d24c94ff5ab6dda63a8d80ee19fd790cb97973328

Observation 088994ec-cb3d-456d-b9ef-65374c8cbcf6 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:05.755965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.475898Z digest=sha256:95fdd978c8f65a086db8671c4c367dfb7620b645a2811a64ded5246ac63c1c7f

Observation 90158bf8-053b-4ceb-a438-182219806833 · outbound

This paper cites aha moment.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start aha moment

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.620532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.555762Z digest=sha256:2c4e79d4e4bf6c64363dd95af868e0cd349f9f9a7b7eb6ccfe9f0b61caa2600e

Observation 676068da-e025-4977-bc9c-ab3d42fc9fcf · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:05.428673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.620141Z digest=sha256:cba4658d36b69dd9d869262e1cd1f72cb57a05d3856c67740b4bccb4a5fd7f63

Observation cdf8577b-67cb-4aa8-834b-6dc6274b82dd · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:05.234616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.674379Z digest=sha256:683b3e67c138b93f83b73f73ff938cadb7a30c97553df49f5987835ece103e16

Observation e1fe6f8b-77f7-46de-bf65-68670cf654c2 · outbound

This paper cites Given: The sum of angle B and angle D is 100◦.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Given: The sum of angle B and angle D is 100◦

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.011044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.740041Z digest=sha256:4c22b10c7ebd1b5467b5f4446fcbd8217bf94f465584069361374e1c7d3249ca

Observation 279d766e-1108-4a94-958d-2a37977303bb · outbound

This paper cites • Diameter BE of circle O means that BE is a straight line passing through the center of the circle.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • Diameter BE of circle O means that BE is a straight line passing through the center of the circle

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.789384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.818951Z digest=sha256:fc8325df84acb980fd7b6bf4df9d242fb87b14a2625d5c5db3f4ef5155563b3b

Observation 3247c9e2-139e-42f3-8d06-0588b65e10cf · outbound

This paper cites Therefore, ∠BAD + ∠BCD = 180◦.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Therefore, ∠BAD + ∠BCD = 180◦

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.576955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:12:59.913549Z digest=sha256:1b03ba3d21eea09e583f7abdf07c339f15874b4bc78beda900fa459c4fb2bd05

Observation d028d651-5586-4630-81be-5ad67db31a00 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:04.430067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.015418Z digest=sha256:13e253ecd5b0c31334ad420fbaef2759ddf15b8b5fe9a73031dc09d102ce4bf3

Observation 5ce9f604-ede0-41a4-a26a-d53905b81d6b · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:04.266127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.095458Z digest=sha256:e67c72406ef13d52d7c77e3272361c2bf778f6c87d4bc7e274c5049da7e31002

Observation f6db60fc-8d5a-4431-b285-6d7689142a1a · outbound

This paper cites The sum of the angles in triangle ADE is 180◦: ∠DAE + ∠ADE + ∠AED = 180◦, ∠DAE + 90◦ + ∠AED = 180◦, ∠DAE + ∠AED = 90◦.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start The sum of the angles in triangle ADE is 180◦: ∠DAE + ∠ADE + ∠AED = 180◦, ∠DAE + 90◦ + ∠AED = 180◦, ∠DAE + ∠AED = 90◦

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.101635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.163730Z digest=sha256:763da81e1d974f2fe7293033875c108b4fca63024b9d7f0cc06e8c965bd10ece

Observation 90fa580d-205b-4209-bc91-e02ff1c20b66 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:03.931966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.225496Z digest=sha256:aa2ef1ff737431332d79fafb6e17d46e8757ed93ce6d4aa6197060fd10fc4372

Observation 59d518fe-a3c8-463d-9e44-addefcc87cbc · outbound

This paper cites Since ∠DAE cannot be negative, we must re-evaluate the problem.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Since ∠DAE cannot be negative, we must re-evaluate the problem

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.744279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.313043Z digest=sha256:e320c04b9d7acd00644ddbdaac77437f87d2ca39f73442cb4cf1d084b5b3d6cf

Observation aa5d843b-6657-4ddb-8ba7-e222869a0b53 · outbound

This paper cites • Given the perimeter is 30, we can find the length of one side by dividing the perimeter by 3: Side length = 30 3 = 10.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • Given the perimeter is 30, we can find the length of one side by dividing the perimeter by 3: Side length = 30 3 = 10

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.577793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.416991Z digest=sha256:c1483365688b9de65fbece08edbe8421f0f53bee3f51fac6ae0acef40b0c3326

Observation 1ecbf4e4-896e-4681-a351-d1a1eabead5d · outbound

This paper cites • In a 30-60-90 triangle, the ratio of the sides opposite the 30◦, 60◦, and 90◦ angles is 1 : √ 3 : 2.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • In a 30-60-90 triangle, the ratio of the sides opposite the 30◦, 60◦, and 90◦ angles is 1 : √ 3 : 2

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.390783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.514135Z digest=sha256:964901295b630a9e19abb1e93cfd0b3bcc7d83dce2331c6098cc1040a57da312

Observation e47c8db1-473f-4ff7-818d-48affed72784 · outbound

This paper cites • The side opposite the 30◦ angle (which is half the base) is 5 (since the base is 10 and it is bisected).

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • The side opposite the 30◦ angle (which is half the base) is 5 (since the base is 10 and it is bisected)

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.218991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.582130Z digest=sha256:5666aa05f9715374991643e0c6e717fcca978c2bb23ac3b5d0e9783041238546

Observation 653666e5-8c29-49b0-a835-47748bb6a5a9 · outbound

This paper cites • We need to find the length of the altitude h of this triangle.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • We need to find the length of the altitude h of this triangle

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.060938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.679380Z digest=sha256:31d01062cd1588d16c2913294361fbce8acbf8441249bd95eb8ca6fd59e01d5b

Observation 2d11caa4-2547-4f9a-ac3b-84425752d182 · outbound

This paper cites • Let the side length of the triangle be s.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • Let the side length of the triangle be s

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.880717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.755861Z digest=sha256:027c898932b094d0c2c55935716a7ddf24646a376d7922202f600a83fea11429

Observation 08494f91-4ab0-4f7c-844e-1574d2307636 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:02.641387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.845200Z digest=sha256:98303ef4f7f76edd56c4cdf3880831a8383832068cf5b0a27bb978a6f368e8bb

Observation ab35689a-28df-41bb-9c68-51d862c34f5b · outbound

This paper cites • In an equilateral triangle, the altitude bisects the base, creating two 30-60-90 right triangles.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start • In an equilateral triangle, the altitude bisects the base, creating two 30-60-90 right triangles

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.413427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:00.932690Z digest=sha256:37c649c8aaa2f3bad984adfd9d774b2ef1b89aad407d3fbf87e0067730d0b834

Observation 7e1946b5-9ce7-45de-9b66-762df84831bd · outbound

This paper cites 5 √ 3 19.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start 5 √ 3 19

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.225475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:13:01.038049Z digest=sha256:cdeeddf9d05499a0e993f2a0ee426c472543bb69e86eaf6d97f1403ada3c1c18

Pith citing papers

Observation d4ad2234-a99e-4036-ad53-08bfec05eca6 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:16.101741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:16.101741Z digest=sha256:e15cc08347ca054e25fa957a5fdaabd70c42945f63d282123865617dcf68ecec

Observation 25e3e49c-c7c5-4e80-a9bb-b7981a14a940 · inbound

ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design cites this paper.

ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.070150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:00:46.618841Z digest=sha256:775cff1beb22dce4a4386abba7350f65fa9a9c38e16fb4e17e08c8d420d81d96

Observation 87ffdd63-c27e-425b-ae71-19dfe4e4e665 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.951658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:0e92631e7ab3c9c33755b02c5763471bad57c1add7ff48b96fe71846bddac801

Observation 2fce9b40-570a-449e-95f3-80cf80901a00 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.492818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:b3a4a14ab077534d820d60464b673c9f522c86cd8280d785b644a025be137c2b

Observation 2e502e2f-c1b0-4b06-9948-b3e3d3a3561b · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.396480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:9a20d96a23f94f619b5c4d195825c1d6d5732d06f8ae5d27e6fb55bbbaa20888

Observation 7c2c75b8-0e29-456b-95e2-f9441410b598 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.698390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:477f15b40db3748f19c058ff898c754b0fcb6b073b2647cfaa8ce2c7d3864727

Observation d2cfd117-e6e7-4385-82b0-78f54c0c56b1 · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:44.686114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:f90a2f7228602d7ca4952e3e82fc20b7894ea03ab854edc8693563e7d489f6cd

Observation d04b7a35-ab91-47a5-a8a3-3928e4eed17d · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.734339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:2feff43c2b2e0eb970bc73f37b2d62cf44fd7a1616bce051308c015e10d6a5a5

Observation 3d65458b-09dc-4312-b131-7cf8af863332 · inbound

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning cites this paper.

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:14:18.868098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T06:12:49.839459Z digest=sha256:540412b1c14d24088c244efc55457598d9e6b5953ba1630958e89d761710e70c

Observation c8eb2883-164b-4563-af06-d99077909789 · inbound

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach cites this paper.

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:47:07.116263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:47:07.116263Z digest=sha256:8b81cbfc773e1079fee027320ef86b9192a2c568214bfc0871f6a973adb1c4e9