Pith. sign in

Paper Citation Record · LEDGER

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

As of 28 July 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2606.22305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22305 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T11:08:54.744334Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:27:07.375901Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91b2f34d-7daf-430c-9466-6102a08dccf3 · outbound

This paper cites On-policy distillation of language models: Learn- ing from self-generated mistakes.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning On-policy distillation of language models: Learn- ing from self-generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:a2006827d0eb1fb862c958be591e30bd293dfa1eb7822f74144e8e3bdf2ce5af

Observation 82051729-61ac-402d-8a6c-c3e22d4fd2ac · outbound

This paper cites Olmo 3.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Olmo 3

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.400005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:3e55e69ac1d18a47917573c6d2778cf4029df5163881d2978ad2010f62d11b75

Observation 1855a53d-508d-4cb7-a8e6-71005afe7b88 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:dfa7de36ee7836b61c4ca24b1b5b76470beea73b6f064bbe370d43315abd0e34

Observation e49585af-c3dd-4990-aecc-91e720eb224b · outbound

This paper cites Curriculum re- inforcement learning via constrained optimal transport.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Curriculum re- inforcement learning via constrained optimal transport

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:d6ed27686de9fe50064e813b5c2de49296d8e17af10896c1131992928aaca3bc

Observation bdbde187-cfcd-4d5d-8c35-b9fed098fab3 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:08691a1b5fefb58c8795ad8814f7ba7dfa30d9ed9674002ff637592ae5d42133

Observation 42d083ec-039c-441c-8ed1-b3f60094faac · outbound

This paper cites Numinamath.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Numinamath

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:3ae205277e6364a893ad473a145ff52af3ad80b3f061c9dd823e1fb1119d16e6

Observation 4b8240bf-6c8e-4654-9049-3a15ac91ceca · outbound

This paper cites Let’s verify step by step.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Let’s verify step by step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:42f75702f7debd9566726d4c992df1f8e8485ed80caaea1580c130291467157a

Observation b696fb04-9c38-43b2-8f51-dc93eec3312f · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.403635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:1636fc2f9173f157610e490025e77b0c9c9e7a6e4584a24f5860e964bfe08011

Observation 9a7e96b4-03d0-4a8d-bf88-d2c65942ea5b · outbound

This paper cites Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.400353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:e4ee01e3abeedaf4404d99daf534faf4c4d623333a51ad4f4fb394209ea3a83a

Observation ed0b958e-6eca-47df-b731-cd9d37ecdfcf · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:d95dbffc58b8779e1e9c20f2da904062eedfc4c5a07fb6c1f9fe410de5fa4442

Observation 4cd62d71-ac62-4d96-9750-e71715369cf8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.397485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:04836bdbd16488482d70ab533ab7e41ae0a97731377663fc0ac819a6d5cc516e

Observation 0dd2c6c3-bfa0-458b-a2dd-e8716cd29506 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T11:09:23.070143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:68cc911f2031cd9ba977ce0da6d0898d562711b25f7d16b399a5ead5172a3345

Observation f74450ee-1775-45bd-8c4e-53d7719c070d · outbound

This paper cites Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay.arXiv preprint arXiv:2506.05316.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay.arXiv preprint arXiv:2506.05316

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-26T11:09:23.065588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:96590397aa869ee04a1828409d69dfaa969f81cab1b2e143835a98df9954486d

Observation cba722a3-7446-42db-bf61-a592c172eb98 · outbound

This paper cites Independent skill transfer for deep reinforcement learning.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Independent skill transfer for deep reinforcement learning

Reference 20

Resolution
verified exact
doi, observed 2026-06-26T11:09:23.066975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:a3cafa270c18a38a11cb34098852ab192ae1563f5ec72def605b6a135cc46c56

Observation 54ff654d-9609-488a-838c-d0dfce81a64c · outbound

This paper cites Mind in society: The development of higher psychological processes, vol- ume 86.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Mind in society: The development of higher psychological processes, vol- ume 86

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:9cda482f01792a2b29a151db7e36c7a1b0f0a6a98f943e8f70670375a83bf0cc

Observation 85c28287-918a-4480-84a6-1f7c78f5f19f · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:4e233c335e6905a191c7eedde99eeea70bd975c6b161d0b3601657eb6bd73dc7

Observation b9fd20ed-4c50-48a2-95c2-ffa5f5e8fa84 · outbound

This paper cites American invitational mathematics examination (aime) 2024, 2024.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning American invitational mathematics examination (aime) 2024, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:0bdada92cfb13cb917be63be77d273024134021855ff961b3ef1047a471a4e27

Observation 8c8c3a20-4754-4bb1-ada3-dab4d9391caa · outbound

This paper cites American invitational mathematics examination (aime) 2025, 2025.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning American invitational mathematics examination (aime) 2025, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T11:08:54.744334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:0e8ce64250b3677e00264cf8be85409e52b81ae85c36d5f5b91a7728a3e30165

Observation ec26ea55-50c2-4794-a6e8-8ef5612c9428 · outbound

This paper cites Group Sequence Policy Optimization.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 31

Resolution
malformed identifier
local_arxiv, observed 2026-06-26T11:09:23.064873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:85d812e6e0a9e2ae02e006dea74b80ce48cff9db506451308af1981772abc6bb

Pith citing papers

Observation e54299af-ac5a-4999-9d10-86526f99c937 · inbound

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning cites this paper.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.377916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:0c47fac4148e8c4f7b1d771ac2fd44bf95a762be2167166d0ad441f165f99892