Pith. sign in

Paper Citation Record · LEDGER

Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2505.13026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13026 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:32.507410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:27:26.462219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43b5554e-0bf3-4687-b7f5-aed6cf3c50d0 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.472508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:a65a76089c49e9ad2d67b448367493e796bcf501506ff8e2c5c0c7bf7817ad9e

Observation 6ae68cf3-9cc2-46be-b62b-fa959f8e26cb · inbound

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization cites this paper.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:56.993587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:ff332cf659d3e264b3f628fcd78bc2ef78369087b15c3c1662eefd83ac4dd5e9

Observation 994317d1-da8b-437a-8e2f-ff8730d29f96 · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:56:54.953338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:c765a048321d4a66054b7008b0b3e1a7cc2d1c835a1ac0f18256b5689681c75f

Observation 1d64f13b-c276-4196-aa2a-14bdc9bec298 · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:32.507410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:32.507410Z digest=sha256:5efc9757299012a7fc5efc3f71a065fef93adf50524bb07e601339345c3bcfb1

Observation 96c74d77-4f08-43bf-b901-c92e189d6bbc · inbound

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning cites this paper.

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:14.415252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:44:34.968571Z digest=sha256:0d1780206edd13ef1c1f4122639f196cc6ef38f5c143a432e405ccd4c608444e

Observation 5b9a0335-f9cf-4929-8a2d-e56052ccac1c · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:17.494889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T12:57:24.822525Z digest=sha256:d2c7397a3525f64bbb0652972942e3df7545a325e772b16e1b636407eaaefded

Observation d1134a14-495b-415f-8224-ae2b5b3a2183 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:20:55.323246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:51:29.021423Z digest=sha256:cfd195b1e492312ff936558246b47f989efa037b977d6a399e29da366f11536d

Observation caf1d41f-2ff2-403c-ae32-e47e3c03e4d7 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:17:59.712736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T21:14:22.718769Z digest=sha256:0a7d14657867d0a4e50fd93e915736beb54df977748b4a006d582a8d13b5c003

Observation 0028337c-5869-44ff-a2b4-d846267ce88b · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.010585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:28f54fd98a2feb7a4aa4ceec36707a4f779c20729159fa10e0069d804e5ae018

Observation 78f0c023-f494-4c08-a0cd-acbfa6d5ac05 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.520517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:92cf0f548baa4ef94e5028e4deccee1426f395e43d58a0fa7daac3dfabf8a947

Observation 81be6abc-94e2-4b28-8412-c180b2b289e9 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.463643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:f8fe3873a879f6bc522407ee5e96294a807f43ba1f0fff33b23273bcab1fbff2