Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:43:14.663362Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2507.07562.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:43:14.663362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T00:14:04.589912Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 89f1901c-ae94-490e-954f-677cf57d7bd4 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb82a37-d36b-4c55-8e65-504b2137b1ba · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7118bfb-52ae-451f-9256-f4511780ad8c · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Arcee’s mergekit: A toolkit for merging large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 932dcfa6-612f-4c2d-8d19-603af71bc7d9 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ada49c-25ff-477c-8b05-69225c260b51 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53ea69e-b35b-4a47-9013-8d1695fc6b97 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs OpenAI o1 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715a276c-f616-4c4e-a226-32d82389387f · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30494fb6-9a92-41d0-854e-af82ae08b37c · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs s1: Simple test-time scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607f44a6-7c83-4c6a-958e-4ea33e61a483 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0947f061-f8f4-4fd4-8074-0c15a6b01239 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs HybridFlow: A Flexible and Efficient RLHF Framework
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042fe905-15a9-48d6-afaf-1cd6fcb6b17b · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4252390-13a8-40bd-94b3-5c13a46eaf1e · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3f7eaa-e072-4070-a191-14f807f18d2d · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7de96f-9c1e-4d68-87df-a3574142399b · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00d0a4b-b544-4070-b08a-39500efd0c46 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5f3f58-337f-41bf-9cac-c7137806a94d · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6993364a-f058-445a-91fb-b92c2850bc60 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f15e2a6-df26-4486-8f23-3ba1178ff97a · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Improve Vision Language Model Chain-of-thought Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bb2bd0-6a99-4730-a655-82aeb47e628b · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da25105-e8bd-433c-94ea-55230efbd11f · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Reason-rft: Reinforcement fine-tuning for visual reasoning
Reference 1985
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a0d5fc-8fa4-4230-8d97-50b9bb59abc6 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24224149-d2cf-4f44-9874-6ae319720223 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fea811-4015-426c-b542-f2c90b4fdc51 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565eaac8-595a-41a9-8e04-6a1bf00cc9b3 · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Unlocking the potential of difficulty prior in rl-based multimodal reasoning
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2b2560b-81c8-414e-9bd5-0cbebde45c4d · outbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4220cfe6-f0bf-41ee-88f0-de5c6f8eacec · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fd3e50a-a2b5-42db-8bf4-dda52bf77160 · inbound
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd0ea6f6-0c41-48e7-8da5-e4f535269f7b · inbound
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.