Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T11:08:54.744334Z
Paper Citation Record · LEDGER
As of 28 July 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2606.22305.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T11:08:54.744334Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T14:27:07.375901Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 91b2f34d-7daf-430c-9466-6102a08dccf3 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning On-policy distillation of language models: Learn- ing from self-generated mistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82051729-61ac-402d-8a6c-c3e22d4fd2ac · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Olmo 3
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1855a53d-508d-4cb7-a8e6-71005afe7b88 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49585af-c3dd-4990-aecc-91e720eb224b · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Curriculum re- inforcement learning via constrained optimal transport
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbde187-cfcd-4d5d-8c35-b9fed098fab3 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d083ec-039c-441c-8ed1-b3f60094faac · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Numinamath
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8240bf-6c8e-4654-9049-3a15ac91ceca · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Let’s verify step by step
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b696fb04-9c38-43b2-8f51-dc93eec3312f · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9a7e96b4-03d0-4a8d-bf88-d2c65942ea5b · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ed0b958e-6eca-47df-b731-cd9d37ecdfcf · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd62d71-ac62-4d96-9750-e71715369cf8 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0dd2c6c3-bfa0-458b-a2dd-e8716cd29506 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation f74450ee-1775-45bd-8c4e-53d7719c070d · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay.arXiv preprint arXiv:2506.05316
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation cba722a3-7446-42db-bf61-a592c172eb98 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Independent skill transfer for deep reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 54ff654d-9609-488a-838c-d0dfce81a64c · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Mind in society: The development of higher psychological processes, vol- ume 86
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c28287-918a-4480-84a6-1f7c78f5f19f · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fd20ed-4c50-48a2-95c2-ffa5f5e8fa84 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning American invitational mathematics examination (aime) 2024, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8c3a20-4754-4bb1-ada3-dab4d9391caa · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning American invitational mathematics examination (aime) 2025, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec26ea55-50c2-4794-a6e8-8ef5612c9428 · outbound
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Group Sequence Policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation e54299af-ac5a-4999-9d10-86526f99c937 · inbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.