Pith. sign in

Paper Citation Record · LEDGER

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 30 inbound Pith citation observations for arXiv:2506.19767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19767 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.148313Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:15.303331Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation cbe01040-8176-406a-b45b-134de8acc13c · outbound

This paper cites The magnitude of the decrease for each non-target tokenvis proportional to its current probabilityπ θ(v|x,y<t).

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning The magnitude of the decrease for each non-target tokenvis proportional to its current probabilityπ θ(v|x,y<t)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.558327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:30:38.148313Z digest=sha256:bb649488a9e8d89dcf75414232912bfa0a4c2ba4b10164ddd9ed3e2f98112c0a

Observation 19f55063-3ef4-4662-b74f-643205027fcd · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.045826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.045826Z digest=sha256:c8780bb8b4da2891d1654a76635e72c46c15c5399968db74f7134384d9f91514

Observation 7775247c-7ce1-4249-84ee-9ace25ee397e · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.050472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.050472Z digest=sha256:279ee61e199d3480ae7a72fe323e5210b37fb57e4a3b5cfeeb7b6fa50d130d80

Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · outbound

This paper cites Improving RL Exploration for LLM Reasoning through Retrospective Replay.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.055667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.055667Z digest=sha256:eda01da6fca93c24e7935d9fc9e3e77de1f92a4711709cac35aaf9f1c24afe19

Observation 5cf16aa3-009e-4a19-9e08-92cbd05491ec · outbound

This paper cites LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.060670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.060670Z digest=sha256:37468c6c6200014606b31fb015ed7a43cb3aa8c070bf035b9c8065b9768b3f24

Observation 5dc7a3d8-dbd3-47b7-8823-1fc9362047f5 · outbound

This paper cites How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.065791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.065791Z digest=sha256:99a8b6444f51b2d3754cf631cfec0181dd7bfa66f1188bfb3f9d2201f5b1e523

Observation 65877c37-2906-48b1-8829-8017d612a744 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Learning to Reason under Off-Policy Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.070952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.070952Z digest=sha256:e2f6d9ef08dc18178383546ad31949e55218255df149ded6e3b08d742deae70b

Observation c0783268-72c6-4f8e-8038-25ed0fae6a5b · outbound

This paper cites TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.074895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.074895Z digest=sha256:30a933fea81447e20c9f130ba5be700dd08a3ccc1ee7e6caa6c35960d496ed4a

Observation 1e18dc3a-2a3c-4b87-9f02-c016dfab8075 · outbound

This paper cites UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.078831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.078831Z digest=sha256:136ad8bc56cbfb9a22b978d055fd02ea25e7545cbd6824186ee29c73089d3052

Observation ad07a096-01e3-40e2-abf5-7afca31de710 · outbound

This paper cites 5, 22 Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, et al.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 5, 22 Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, et al

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.092173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.092173Z digest=sha256:ec8d304141c194973fdea6b983d8ae2e40c0b91e92926fa1122e7b2771db4b47

Observation aeeb7730-bc28-48cb-8a7d-8a72a80e125d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.100595Z digest=sha256:93c4895b69f175a07e13852cf9882c1387084ab736b6bf879e18f32aa826f503

Observation 84bf90af-a3eb-4d56-8eea-75b646a479d1 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.109768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.109768Z digest=sha256:f69b57c7dd4ddbd132f865866369fcce7efa966ac4f4f1ef7be24784f7737d70

Observation 1e80430f-21f4-4522-9a3c-1321d14f7aeb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.114032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.114032Z digest=sha256:ee1c3a6fe2e27d413f7dc9b34f1a7843389d4861c4cfafc3a3d3934eaecbd93a

Observation 426a2e7a-f3f8-4cec-8b43-72664517a6fa · outbound

This paper cites Process Reinforcement through Implicit Rewards.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Process Reinforcement through Implicit Rewards

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.118345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.118345Z digest=sha256:8c1c7c292d69523ce2339c02f5e1fc73c0298c100a9c2ac84d3adea8a42c6595

Observation d4b4cafe-138b-4da1-b556-255e515f2772 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.122778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.122778Z digest=sha256:74be3e839b3c138a7831f2b06d6f501bbd4e568ecdb6693e837ea120ac0b422d

Observation 7c0839fc-e534-46bd-b906-90277b620ff8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.127553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.127553Z digest=sha256:0c45e9b98f2a4870c14578fc368252248e8726e16f676baa0c0cc1c9d48a5465

Observation f77ef34b-767b-4b0f-8fd2-74e212b73360 · outbound

This paper cites For datasets with limited sample sizes (AIME24 and AMC), we report the avg@32 metric; for the remaining three datasets, we adopt pass@1 as the evaluation criterion.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning For datasets with limited sample sizes (AIME24 and AMC), we report the avg@32 metric; for the remaining three datasets, we adopt pass@1 as the evaluation criterion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.601772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:30:38.136077Z digest=sha256:94d6565696b0479fb0829871f415f369f88efd04a3ce622dae284dc2fa8428a6

Observation 823de888-79d3-4a9f-bb10-161813419857 · outbound

This paper cites <think>\n {thoughts}</think>\n.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning <think>\n {thoughts}</think>\n

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.586601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:30:38.140023Z digest=sha256:888e7ea3ebfac78a87171770ad1354b205f66a6d88047da7880c2ecc043a0e19

Observation 147acd18-44e4-4828-abf0-79a993c358e8 · outbound

This paper cites Google-proof.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Google-proof

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.572625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:30:38.144291Z digest=sha256:4a129f524a88ad69638d86b9264b162be18dd671fa2cfc1f0c4c789b74e5888c

Observation 41f8982b-c858-4a83-8a96-bd5958b82045 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.083143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.083143Z digest=sha256:088130758449b0b36f4c656a7f6f871290440e6b17f35cdbad73a28bcc311198

Observation cf1c6665-0cd0-457f-aa09-037c426ab0b7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.105293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.105293Z digest=sha256:352114001334ddbc77684fdacd053d36a11f073abb6495f22a3421729df83c3d

Observation fe7424f0-86da-4f18-b9f3-6817648067c8 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.096229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.096229Z digest=sha256:bfe446a7d1bd199e3d9cda469c6760025a550ace93ffc8f95dae8567de093803

Observation 28f53627-385a-43b0-84cc-c2ce833bc023 · outbound

This paper cites 17 A.2 Entropy-aware Gradient Clipping.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning 17 A.2 Entropy-aware Gradient Clipping

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:38.616635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:30:38.131784Z digest=sha256:ed147227ffdebc98fb056ff6faff8631354f1b80ac3faa0ec460086f642a13e5

Observation ff13ae6f-403c-4c05-8609-00d8adfe5d6f · outbound

This paper cites Proximal Policy Optimization Algorithms.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Proximal Policy Optimization Algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.087653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.087653Z digest=sha256:d69ec9cdd25d8ba0a809fc3406ca2f7c28fdcf9e4c8e1a41d8ef44579f17ae78

Observation 0f5e7844-bbd9-4cb7-9fa5-b754265c7a04 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.040628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.040628Z digest=sha256:c3711b14a01ac3874656bff11c5c6c3823c61fb638e264b8c41aec85fc72eef0

Pith citing papers

Observation 9ce80c5e-b8d2-4b1e-8cb6-e72c91dd252f · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:23.526899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:f63c37b2c77e9d6b8164850cf64793883d79afd1e218f6eff129646349339b84

Observation 6c3f19e7-c0d6-459f-924c-eb6dc8dd69a3 · inbound

Post-Completion Learning for Language Models cites this paper.

Post-Completion Learning for Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:15.303331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:52:15.303331Z digest=sha256:d538973b90089108afea7eb5af513cc826f5f2234a32c9406d903de644c923d9

Observation 3f2dbca8-2b50-4732-998c-73f84649e9bb · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:56:55.045505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:36a470d18f456a6fe9e741a677985f4fca6bf28c6a0e998df24a41ca55133331

Observation f9961d65-2c5c-4622-923a-f528de674486 · inbound

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning cites this paper.

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:33.169961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:33.169961Z digest=sha256:851a186ce1da4d61db5627b5a578727629132374b8171f52f149f57135cf2b21

Observation e36dad10-0739-4a45-b78b-9b7028f14b24 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:36:53.701052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:7d5ebaeb69889fc8df9faf4131c2bfdabd7fcc9eb98958119313ab37390a6b41

Observation 2d21066e-700b-4b6d-a68f-377c3c099a3a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.405749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:ef815db8ffbe38f50284250ea929cb092142cbf389d2a82801c2b46be8d6a965

Observation 975e4c8d-006c-480c-b21b-c252f9e63794 · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:33.241269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:33.241269Z digest=sha256:f6d0e9afabf7d7074b20062a35324ae72437094cdc7124191b81cb007d089036

Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · inbound

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models cites this paper.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.176602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.176602Z digest=sha256:2754bc3286d0efe0fff8a5d9b15c7b953f6630ff9dc2cd026b2f969aa33e1dba

Observation a44f1f16-1fbc-4dc5-bee6-5a444af496a9 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:26.597119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:26.597119Z digest=sha256:6103d57f750e22a2cc834c68f9b3cdc240051be04b08d57d0757cf8aea52633f

Observation 95a5c8a6-8a44-46ff-b8ce-25da8195f68b · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:38.707015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:38.707015Z digest=sha256:6b2d257cc0e6f4a68178f2cecb7a735a05688bfccf124f6416b696a8a30983f8

Observation 336fcfdd-e5d7-437a-87d2-d788895f128a · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:50.322424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:50.322424Z digest=sha256:e96eb8af3a52caa75fa8abe8b0572174281a42587569f89f9b8dd1b4d1b6bbb7

Observation 50b95321-bffc-40e8-867f-d8c974c23178 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:43:37.638943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:89be6ba0bcd928c0802e510d78b14c998a1ea660290908bf286d57fc98637b92

Observation 5853afe1-a468-491f-b111-8e395ab4a9a5 · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.662347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.662347Z digest=sha256:797cb56536f0fa95c020b64495cd4b919526dea45c5bb594b884d70f541d51e2

Observation b00e54c2-0f98-4160-835d-a032820f5cf6 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:45:37.356079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:5fd887b021998e1a2537f6332f3634e80ed09394ebd0f96140af609695e33056

Observation 587b056d-a5f2-41cc-93fa-1a211b8b1ed3 · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:27.829704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:51c926a4c1c0796b9c3a3c21c0b4b51d1477be402021d1ecfd218f64538e0a98

Observation 3b331609-45b2-472c-b100-ac69b4472109 · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.662477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:176626a9c6522de761c5aa33d8cafc4843c175c7bdd581cb0d7d6f9d2cad4634

Observation 7c4bd75f-8c56-4966-a086-5eedfcbb4d73 · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.011351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:c3ec0c67d0af46f7e76ae7bc2ec2ba560b0a080cad66d4fc04e3f12bef34b2dc

Observation 285b16b3-efde-4ca8-89b8-1e5350cd1871 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:31.724069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:c30a76f5a0a9c00985ec8ce132cd60dbb93714dc422fa93408cc2ecf5b9b39a9

Observation 83613a1d-44e0-467f-8c7c-ef40c35fc3ac · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:07:42.235985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:03b5e539e6922146712fd28432db7a5b69c4f5fbf266c4744f3770481ffd2ecc

Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.645623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1958ede25d3298a68ae84e4063a53a47adc19f2e11eabe04beccc3e18f186dc0

Observation 96a80a11-f43d-46f4-a21d-eb95fc43ec92 · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:25b0af3e5cb7ef0f9adef83a31fd003b68babe54e40f0c598fd23e36199692a4

Observation c60b7a7d-435d-4f00-bc9e-a0ce27f0fe67 · inbound

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training cites this paper.

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.719169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T22:29:05.467252Z digest=sha256:5bb487abdb9389124b290b8c4c73dbaa357fe41e2e80f9c8ae854de192103e0d

Observation 21f8c1db-3f48-42ea-9da4-fb704c60b231 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.496547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:5102905d8caa03aa4c7034cc37f8d96a5603281ca895c812fbb22709ea29ca9e

Observation 967da951-9120-4295-8a7c-68731a45b4d7 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.590285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:8899ee2ac72b478c61df7f5c910aa48109e06c786123cfdbe02fe82c2ab52734

Observation ee13d56a-e9c9-4aa0-9614-253e730673ff · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:56.991869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:98bb82bd24d0a80abecd8817f3e096db7236192bea863263e9aa0345d5af5866

Observation 94116e2e-ed12-4903-acf6-fe4865d5320d · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:58.442808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:7280e67d7dc756990bdfbafaf17968a2d060176665a7d450d09d9a60dfae59ef

Observation 151eb895-fbf2-4d7d-ba72-a03c726fcb31 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:90e1f6e5f96bd59a5aee5f7ce2cfcdda4d4e61b125baf8825ff1e0e6f8098c60

Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.545828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.545828Z digest=sha256:5a828cd4610957997958cabcc18e2620145caf9150c1cc60879a18a180228ecc

Observation b13a36df-6e61-4002-bd65-f7408ccdc670 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:48.101782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:48.101782Z digest=sha256:f9fd5d222c98eecb5880999008c6c9372ee87914ba3a839aba7fa5b7716cf523

Observation 66463944-db0e-4f81-8a74-1a8135b269ac · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.937566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.937566Z digest=sha256:4c7294fb3895142037b510f2cd2f7a6b30054e8125b9c193b6debcd4eb6afd38