Pith. sign in

Paper Citation Record · LEDGER

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2507.07562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07562 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:43:14.663362Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T00:14:04.589912Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89f1901c-ae94-490e-954f-677cf57d7bd4 · outbound

This paper cites Qwen2.5-VL Technical Report.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.071992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.071992Z digest=sha256:f420f476d494559f032d7de25d280a115eed7e27b5ad913884559de1c3bf4c70

Observation 1bb82a37-d36b-4c55-8e65-504b2137b1ba · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.351683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.351683Z digest=sha256:9a49e19ebeccfd931192dd179c0853e904473fb6f3f1c86cc6ddf386698955be

Observation d7118bfb-52ae-451f-9256-f4511780ad8c · outbound

This paper cites Arcee’s mergekit: A toolkit for merging large language models.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Arcee’s mergekit: A toolkit for merging large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:43:15.589359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:43:12.452476Z digest=sha256:a32c5a7e900ab3bd8a81d7ce95ceb4c1a669f9b8b09760b6ac935b927a9ec2cf

Observation 932dcfa6-612f-4c2d-8d19-603af71bc7d9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.533076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.533076Z digest=sha256:52ee7626113b54fa47526e1da23cc639f78489ccd58ab1422dd4481c5683bdfe

Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.605646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.605646Z digest=sha256:d9784e424d7ec9e1ae69395cc1f04e579746decb1557cfa67273197f1914d5f6

Observation 18ada49c-25ff-477c-8b05-69225c260b51 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.666793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.666793Z digest=sha256:ed01dd1cbb8d5c8871e7ea5781dd913dde1aacfcd8ca4b2643d81748354dbacf

Observation f53ea69e-b35b-4a47-9013-8d1695fc6b97 · outbound

This paper cites OpenAI o1 System Card.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.819360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.819360Z digest=sha256:3ec87e45034e080c14d9c16a4ecebe2556ed5801d6a3475a65216ca4f0007a34

Observation 715a276c-f616-4c4e-a226-32d82389387f · outbound

This paper cites AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.895530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.895530Z digest=sha256:adac9636bb9732e37c68bdc93f37e3c4bfb54f5a97b179a2bf1b396984a03810

Observation 30494fb6-9a92-41d0-854e-af82ae08b37c · outbound

This paper cites s1: Simple test-time scaling.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.038090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.038090Z digest=sha256:1f75016e31ca6f565fe84955773de93204cd88cba641d41357ad0ba7263575fd

Observation 607f44a6-7c83-4c6a-958e-4ea33e61a483 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.127934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.127934Z digest=sha256:dd25cf8e0018f0a68b66a4a0ecfe46b1b8602787577d6a10a957a7ab4239c36f

Observation 0947f061-f8f4-4fd4-8074-0c15a6b01239 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs HybridFlow: A Flexible and Efficient RLHF Framework

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.188974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.188974Z digest=sha256:24f242ceb704dbb45d91d69054c649c264c4bc8647feec40f56899b067f2ea99

Observation 042fe905-15a9-48d6-afaf-1cd6fcb6b17b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.310210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.310210Z digest=sha256:c28786a2a11fdf0a5d9c215fd62f540fb6167479e96610910592c6f9b74ed07f

Observation e4252390-13a8-40bd-94b3-5c13a46eaf1e · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.368576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.368576Z digest=sha256:71950e00dea95b11cd13d72ad4663208a29fc9ac4293aa7d613c99f7b96847a5

Observation 1c3f7eaa-e072-4070-a191-14f807f18d2d · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.461715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.461715Z digest=sha256:48e9b2daaf8b99f41ea9fc983558d3aa44d2e69f7af2c578c3704957901bebd4

Observation 8a7de96f-9c1e-4d68-87df-a3574142399b · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.654093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.654093Z digest=sha256:96a132fb79b92406e944db1c88121134e0e3efbfbf15bc3f9880e4421c5ba8f4

Observation f00d0a4b-b544-4070-b08a-39500efd0c46 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.832356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.832356Z digest=sha256:d95b424df36e1633ff69fb65518dceb346a15928a7ede5510892deb69de3d7f6

Observation ee5f3f58-337f-41bf-9cac-c7137806a94d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:14.145096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:14.145096Z digest=sha256:4f167ef64683935ba0cf0f8774da67ece042079e7e3cf9bda9e22fca577389cb

Observation 6993364a-f058-445a-91fb-b92c2850bc60 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:14.386503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:14.386503Z digest=sha256:da274d4db875cd546480997c3e93bedd0322e151ba41e0246ffb3a81eea9e014

Observation 2f15e2a6-df26-4486-8f23-3ba1178ff97a · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Improve Vision Language Model Chain-of-thought Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:14.536799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:14.536799Z digest=sha256:0217cc35d624ac54921ad5a76d88423de4da10b4a7e9b26d1c57091b170b96fe

Observation 66bb2bd0-6a99-4730-a655-82aeb47e628b · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:14.663362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:14.663362Z digest=sha256:c2b2470b09ad142ff00ba29225bcfb30536991306d809e3c44f4d52bc6d9baa3

Observation 2da25105-e8bd-433c-94ea-55230efbd11f · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 1985

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.245905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.245905Z digest=sha256:1166908cc63c9fef2ac3dc15b7f934396f5ac25452ef3944a7ceb0caabe5841f

Observation c6a0d5fc-8fa4-4230-8d97-50b9bb59abc6 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.737770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.737770Z digest=sha256:df70e24d806aef753b9f198804a8498d8d07c4daecaef4bc703a7c6e5ae68ae9

Observation 24224149-d2cf-4f44-9874-6ae319720223 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:13.535549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:13.535549Z digest=sha256:7ecfc93987659a65da53a1586a5ba7742b6c69439bffb32db95c3460dfae4fa4

Observation 49fea811-4015-426c-b542-f2c90b4fdc51 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.963346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.963346Z digest=sha256:521bf775db70b7bc62e9f79796c7d458b91fa957ec799fff287242722cd07526

Observation 565eaac8-595a-41a9-8e04-6a1bf00cc9b3 · outbound

This paper cites Unlocking the potential of difficulty prior in rl-based multimodal reasoning.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Unlocking the potential of difficulty prior in rl-based multimodal reasoning

Reference 2024

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:43:15.339676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:43:12.271902Z digest=sha256:61d9d79e390e93fae0b7eb89d5478f388120cb9b850e1406c276f820cd83c0d4

Observation c2b2560b-81c8-414e-9bd5-0cbebde45c4d · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.147943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.147943Z digest=sha256:1564a9f440e1d4e66552109228ad170265236a27a1030ca7d47f321a4fe320ee

Pith citing papers

Observation 4220cfe6-f0bf-41ee-88f0-de5c6f8eacec · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.815805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:ac81a3a2174274524bde14e0bbc7f5242abd7bbf68bab85a825e1e56c13dc881

Observation 4fd3e50a-a2b5-42db-8bf4-dda52bf77160 · inbound

On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training cites this paper.

On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:48:00.469365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T14:47:52.094353Z digest=sha256:2b1621374382855d95a6f45547fb9c22ad71cc871388cdd1aff3f8e684bdbd80

Observation dd0ea6f6-0c41-48e7-8da5-e4f535269f7b · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.591636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:eec5a2e51b28739133c9ae9d0c2dd0e41f9c46315a55c6567150832c6347acb9