Pith. sign in

Paper Citation Record · LEDGER

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 15 inbound Pith citation observations for arXiv:2507.07451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07451 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:53.026897Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:59.069553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.165966Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b088bdc2-b2a5-4392-a83f-8feaf9ade318 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.580139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:47:52.741698Z digest=sha256:3074e58730a3165998359d1ffc4b99288981de85daa59fce112f67c76e124140

Observation 73a3f962-c41a-4bd2-91c6-7e2c8e929205 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.863138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.863138Z digest=sha256:775f06a000564c654bb6405ae6b0e9b7cd0176f80da805e1dded7975cbf168d6

Observation e44144e7-bbef-4dc3-abaf-78aa41bf1c3f · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.565275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:47:52.963492Z digest=sha256:679b7bead49c3c697aa11330786ec98c23b9a20eb9adc99fcb3ffc60d6328abe

Observation b559eaa4-2bff-45fc-9c63-92603b92741b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.967896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.967896Z digest=sha256:f643421695db8b6085a503d0a8f730721f4a93a062c79f464e177f4fee2f373d

Observation 52e57824-a06d-481f-b04a-5e8790af495b · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.972497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.972497Z digest=sha256:222cf34432b198d481d5175af05190ff670593370ed20dc70eede1b713d5c378

Observation 474fbb0d-2761-4a12-a0ba-c3bdcad7ada2 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.976774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.976774Z digest=sha256:03663ee53835ff1486782681378e854b20a051a38d2790f3dc8b7309700784ef

Observation 9507fb39-46cf-407e-afec-818af53a1e7e · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.550737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:47:52.981900Z digest=sha256:82a656ba454e6726fcd82fcd512a14680442fee3a2a790aab08f9fb91fe6fd93

Observation 9477c727-1f68-498f-820c-9995f14a3279 · outbound

This paper cites Openai o1 system card.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Openai o1 system card

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:47:53.534839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:47:52.986066Z digest=sha256:88b7504094987c85ee876151d0152890e4da9c92ff19fa944cd2222010362943

Observation 4ae26289-3377-4e46-8f47-369694799228 · outbound

This paper cites Prioritized Experience Replay.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Prioritized Experience Replay

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.990150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.990150Z digest=sha256:e200f5b057133afaf9b6ad51cbae42b0f610b3a1cbb9f77533e74e3329f83a80

Observation 73bca3c6-127e-4ed5-8307-793c812424b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.994487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.994487Z digest=sha256:fec3eaf09488a932f5c457d5d6bc767b4844805fa693fa5e42865254d03af910

Observation 953f5c33-f686-48ba-b15d-14589505cdf6 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.998807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.998807Z digest=sha256:28eba317e2bf495e7635e815fdccf051c9849b01651eeb03f029d35c1a7c6816

Observation 89149ea1-74f9-4a6e-b5a6-fc3827d7b7b3 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.003183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.003183Z digest=sha256:0d51d7f8843437431ece06ddf849a048a7a9d35a0159bf4d5e68b64a6d3cf992

Observation 70b477f0-1efe-40e1-acb5-a0f6cbcc92b2 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.007802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.007802Z digest=sha256:0e7ddf21de9f1c56793cf5b9cd8a5075671669711f0c5d57e307094a5866d0e6

Observation d01f2037-0923-4a59-94f8-19a5ee13f10d · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Learning to Reason under Off-Policy Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.012363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.012363Z digest=sha256:dd9e60d92c45a3355b2cfe33858a766ef07e79560770fff25adc6d695968e549

Observation a1c3d17e-85a6-4121-b93c-b096f2e75002 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.017126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.017126Z digest=sha256:7283dadd5f1977307b4bd38673453dffd2e213e1f1d088efd45d129fb1acf4f9

Observation 99272396-9262-43cf-b742-4022c013c4f1 · outbound

This paper cites Qwen3 Technical Report.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.022014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.022014Z digest=sha256:8e5fd092c4f65036c06a72cadeee3889cfc6795d81731f92b15122eae6fb182e

Observation 753b1d4b-3e9c-4d5f-b4c2-96ad20b0e56f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.026897Z digest=sha256:80d7d2912845f986150a275c934be579b13bfff7928c30c26b169f5d34b5279a

Pith citing papers

Observation dac9cfc3-8c42-440f-ae91-ff827c37b32c · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:59.069553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:59.069553Z digest=sha256:8cafdf4f0c7fb72b9ab1a39f3a6d0fd55fbe1337739ab84b830fa871b0be701a

Observation d7f21a67-f54a-4b20-ac92-4d3c06536a5d · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.748293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:d2ffbfc8387eb0cdeec91991e2a07b6430a6daca2364e6ff2b5d9b335157b2fe

Observation 2092eeef-c21b-4e07-9145-e23e6f9a9eaa · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.603133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:270a5cec7adb68fbb486dbc456eeed7c33ba08cae941e8e5a9755df8c5e497f6

Observation a8b9d20c-620a-4815-9db1-9c01ea163064 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:55.227343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:cad5ebc41046ebf888f927514b5428e4b059bcd7d6737e936375aadf687530f7

Observation ac6c1eb1-842d-4db9-8521-9f10ca992b96 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.655426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:283f4edb2ecf4216d7184e1c685d23c9f9e802aba4301fb5f8c0fdc0b8ef60c8

Observation c4d71b5c-3edd-49e8-8ea1-e94e8cbd8310 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.343127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:9dd8efd23f8044030e1477999d6ce02ad185daed467c5910d39a19329745a338

Observation 7e06d7af-d662-47fa-95f4-a09215f3a59b · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.002327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:7a476177fc3cbb57174016c235e52eb4d51c3974a115b0d1f408870b841c2377

Observation ddd15fc1-9f0d-44a3-8ea3-1575dfdde3e1 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.007463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:cbeadd3eaa697dc54443edc8e6c80430c20dd90cd47f6163deb1a42dc21ac17c

Observation ceee3d62-14e9-492e-b822-b47524bb934e · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.024146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:998aad77cad6a218d378a33bb651f76fa3c2246170b22304def427a25d6a8926

Observation 08963af0-1e52-4b44-8a39-67a91b27948b · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 261

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.167556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:6ed3791d10460eb9be30d09ddf150367dd698f089c9e3c91f0e41e554d753d74

Observation cd42a4e4-5c79-42c5-b36b-22bfd527b87d · inbound

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training cites this paper.

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.451888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:19:00.012926Z digest=sha256:ba2c5f4067439970ff9a7307d3787ffa1d8a30aab428b95761da5e4b9ef69855

Observation 6e9c4b3e-42fc-4ebd-b38b-3d63fb707d2e · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.299729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:52:55.368187Z digest=sha256:a9fcad8eb78d31d68d20ad43a9478b2ca55d0313273c58dc2126fc61bab32b18

Observation e9e37322-2c65-4c7c-9231-098e61e842ed · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:16.230452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:16.230452Z digest=sha256:9600828ff88f567356efc965cf28d377233197a79640682478daee6e87f26e44

Observation 501b1519-0b51-4cc8-9d05-ace24da922c5 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:873d12478772ead23499859c1189d4be4f5e948a1788c8a0d863ea771bc03f09

Observation e769e426-7214-4d05-bd84-4105242bdd2e · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.013207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.013207Z digest=sha256:f0e0665d6ac8a0f61fbc7599f90a4bd10a148c14838ff03f8ebe32ff4b3acf8f