Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:53.026897Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 15 inbound Pith citation observations for arXiv:2507.07451.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:53.026897Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:59.069553Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.165966Z
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b088bdc2-b2a5-4392-a83f-8feaf9ade318 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73a3f962-c41a-4bd2-91c6-7e2c8e929205 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44144e7-bbef-4dc3-abaf-78aa41bf1c3f · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b559eaa4-2bff-45fc-9c63-92603b92741b · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e57824-a06d-481f-b04a-5e8790af495b · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474fbb0d-2761-4a12-a0ba-c3bdcad7ada2 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9507fb39-46cf-407e-afec-818af53a1e7e · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9477c727-1f68-498f-820c-9995f14a3279 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Openai o1 system card
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ae26289-3377-4e46-8f47-369694799228 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Prioritized Experience Replay
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bca3c6-127e-4ed5-8307-793c812424b6 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953f5c33-f686-48ba-b15d-14589505cdf6 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89149ea1-74f9-4a6e-b5a6-fc3827d7b7b3 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b477f0-1efe-40e1-acb5-a0f6cbcc92b2 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d01f2037-0923-4a59-94f8-19a5ee13f10d · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Learning to Reason under Off-Policy Guidance
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c3d17e-85a6-4121-b93c-b096f2e75002 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99272396-9262-43cf-b742-4022c013c4f1 · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753b1d4b-3e9c-4d5f-b4c2-96ad20b0e56f · outbound
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac9cfc3-8c42-440f-ae91-ff827c37b32c · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f21a67-f54a-4b20-ac92-4d3c06536a5d · inbound
EasyVideoR1: Easier RL for Video Understanding RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2092eeef-c21b-4e07-9145-e23e6f9a9eaa · inbound
Near-Future Policy Optimization RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8b9d20c-620a-4815-9db1-9c01ea163064 · inbound
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac6c1eb1-842d-4db9-8521-9f10ca992b96 · inbound
Trust Region On-Policy Distillation RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 275
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4d71b5c-3edd-49e8-8ea1-e94e8cbd8310 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e06d7af-d662-47fa-95f4-a09215f3a59b · inbound
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddd15fc1-9f0d-44a3-8ea3-1575dfdde3e1 · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceee3d62-14e9-492e-b822-b47524bb934e · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08963af0-1e52-4b44-8a39-67a91b27948b · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 261
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd42a4e4-5c79-42c5-b36b-22bfd527b87d · inbound
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e9c4b3e-42fc-4ebd-b38b-3d63fb707d2e · inbound
Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9e37322-2c65-4c7c-9231-098e61e842ed · inbound
Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 501b1519-0b51-4cc8-9d05-ace24da922c5 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e769e426-7214-4d05-bd84-4105242bdd2e · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.