Pith. sign in

Paper Citation Record · LEDGER

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2506.19767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19767 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:48.101782Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ce80c5e-b8d2-4b1e-8cb6-e72c91dd252f · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:23.526899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:1bf6b29e1f3e7fd5671072f61086983e5fa01e51dd04198604a21697886213d6

Observation 3f2dbca8-2b50-4732-998c-73f84649e9bb · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:56:55.045505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:44de5f1872c0ab78960c0eadb41fb233f05d05f778ed7fb2c58f85c2629dd053

Observation f9961d65-2c5c-4622-923a-f528de674486 · inbound

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning cites this paper.

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:33.169961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:33.169961Z digest=sha256:cd019cc6ee0914a17dd308bda660a72e367998c13143513800bb5ca009937eeb

Observation e36dad10-0739-4a45-b78b-9b7028f14b24 · inbound

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration cites this paper.

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:36:53.701052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T22:33:01.074518Z digest=sha256:74657bf52f77e257278861dfde72efaf9f39c0092874837ff491d9fe5b6b0f67

Observation 2d21066e-700b-4b6d-a68f-377c3c099a3a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.405749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:de53ee9d2f493b57f40ed7d0f6422b172c15374dec066e3cd4288e05913ca0ea

Observation 975e4c8d-006c-480c-b21b-c252f9e63794 · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:33.241269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:33.241269Z digest=sha256:78a00b03ffbb8055804c49e6a66709329d8215c9d6c25d2213f453aad21db523

Observation a44f1f16-1fbc-4dc5-bee6-5a444af496a9 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:26.597119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:26.597119Z digest=sha256:e36296ea7288532c25b04956be23a938c50e9ab54d2e63ad983b4d4e55b89aae

Observation 95a5c8a6-8a44-46ff-b8ce-25da8195f68b · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:38.707015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:38.707015Z digest=sha256:ee2c1e1c5e7693487e90be4427aaa25054cf9ced798337c829438e5abaa05cb7

Observation 336fcfdd-e5d7-437a-87d2-d788895f128a · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:50.322424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:50.322424Z digest=sha256:54fcbf088479bff82b6b2cff213d8dcb06d8cf930378331b6f37775d7559c03b

Observation 50b95321-bffc-40e8-867f-d8c974c23178 · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:43:37.638943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:16fc0aa8a6d1020fc080ab1309a023f818083722f2a551fa0841c86e28558a2d

Observation 5853afe1-a468-491f-b111-8e395ab4a9a5 · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.662347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.662347Z digest=sha256:2a39f5b9ce6bfc161174352c7790fb2cd9665cd7335f84735bf514c58ce29b5b

Observation b00e54c2-0f98-4160-835d-a032820f5cf6 · inbound

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings cites this paper.

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:45:37.356079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:44:50.752937Z digest=sha256:b529fa5699e915dbdbfb82acf3849a26a086ba6dac055f8036d6a9750bf6e8ca

Observation 587b056d-a5f2-41cc-93fa-1a211b8b1ed3 · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:27.829704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:e88251be2a92da3677a487384dc0efa330e61a6a4c71ccb65808dc3f5af0dfd2

Observation 3b331609-45b2-472c-b100-ac69b4472109 · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.662477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:e96de8e23cf75ed55a6bfef1eb65029910b34f12ebab8b067e8fdec77cf6a8f6

Observation 7c4bd75f-8c56-4966-a086-5eedfcbb4d73 · inbound

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors cites this paper.

Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.011351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:05:51.423427Z digest=sha256:334f082fcafcb6f698f1af330cb046365ee0da658e2a44aea1a726db6a8d8cfa

Observation 285b16b3-efde-4ca8-89b8-1e5350cd1871 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:31.724069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:448504668ab61a6d5eb6e9ada377198da44fec6a039c2912d6dcb7842f6cd8f1

Observation 83613a1d-44e0-467f-8c7c-ef40c35fc3ac · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:07:42.235985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:9e5a45562bbe9311aa88711522c6759b8fe4fe842a387e5755a3544f3d3d4a57

Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.645623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3de5f9c3760c00ea2ddb27fdeba6d053404a33adaff099ce7ca09112dc5fac29

Observation 96a80a11-f43d-46f4-a21d-eb95fc43ec92 · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:b9f5df8384c22970402a455fa1fffc316c010d57e6faa25345fd982487b72b1e

Observation c60b7a7d-435d-4f00-bc9e-a0ce27f0fe67 · inbound

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training cites this paper.

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.719169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T22:29:05.467252Z digest=sha256:bee87565263d8a5ce5d9b9922170d2b903bf90df70f535fb244dad1ec281de6b

Observation 21f8c1db-3f48-42ea-9da4-fb704c60b231 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.496547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:fa09b81a2281b497f286692e509898c7d7c58fef9f59760141bfd64825e61df1

Observation 967da951-9120-4295-8a7c-68731a45b4d7 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.590285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:4a49592fbfce21761ff64634e3c95962e3c144529f2b11a4552edfd9b932bb40

Observation ee13d56a-e9c9-4aa0-9614-253e730673ff · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:56.991869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:0bc699bbaee80448b96c3f2334a1edf6e1dc1b5de938d92f0d02673b6e6ec495

Observation 94116e2e-ed12-4903-acf6-fe4865d5320d · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:58.442808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:02607b8df3f18fa3c94908490eca25063539bb9fd99acc783b8a03f702c33ca6

Observation 151eb895-fbf2-4d7d-ba72-a03c726fcb31 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:82dae9eb61c8730b600ea2acdc88797ef0830dfbe65d3c97c847600beaca27f9

Observation e7b7907b-ccc5-4b9b-be3f-b7943ba9f25f · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:28.545828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:28.545828Z digest=sha256:93174a58fb529959ea00d55185efe153faa0ac7ded49f035269cc192fad76662

Observation b13a36df-6e61-4002-bd65-f7408ccdc670 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:48.101782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:48.101782Z digest=sha256:81190d40a0029aedaff53b42d754c05260fd0eb4423943db1490c360d0ed9f65