Pith. sign in

Paper Citation Record · LEDGER

Improving RL Exploration for LLM Reasoning through Retrospective Replay

As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 6 inbound Pith citation observations for arXiv:2504.14363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14363 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:29.001062Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.055667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.189087Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 888524d1-e700-4d13-b272-6e7915a7e389 · outbound

This paper cites IEEE Signal Processing Magazine 34(6), 26–38 (2017).

Improving RL Exploration for LLM Reasoning through Retrospective Replay IEEE Signal Processing Magazine 34(6), 26–38 (2017)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.505944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.505944Z digest=sha256:a1bd686b9a93129c57c50e6b42fe61d3ad1abdd33a5dd7719badcb4a2b597987

Observation dae19bd0-8887-42d9-9b17-6e1071ebc1e6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.575439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.575439Z digest=sha256:ac9a617a84f72e8436be329a4c9583a2892046fb317e52a77f55793b94d7c4d2

Observation c90a7971-3e79-4e58-bc1b-850676693ab8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.666657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.666657Z digest=sha256:a6cd60cf9a8420a56a5900ec91506e0045a8492699176c5d19f8dc10989e4662

Observation 9c39bf08-9c78-4d61-9ebe-bcffa6d48bb5 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.671172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.671172Z digest=sha256:6369f3650c8e0ea2d282b817b0a76146cfca96917afd46ed3babb0b22b58d41a

Observation d20b8933-9a52-46b1-bfc6-e781c9329f0a · outbound

This paper cites DocFusion: A Unified Framework for Document Parsing Tasks.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DocFusion: A Unified Framework for Document Parsing Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.675801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.675801Z digest=sha256:7cc0921bb01e6ffae47b2b4bc9ed6cd373171863d6a9fe0e81ab8f061dcc13e0

Observation 42bc350b-17e5-423b-92a7-b384c0860eac · outbound

This paper cites PanGu-Coder: Program Synthesis with Function-Level Language Modeling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay PanGu-Coder: Program Synthesis with Function-Level Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.679663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.679663Z digest=sha256:c30a3757b0ad8ae9bebc76ad2f328558c8c9e6825081426ea4ac503fa53e398d

Observation ad4659b5-0974-46d4-a975-452382807278 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.684280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.684280Z digest=sha256:dff75dff0e25018a8ac04e142a1584b9b45317a9e5975f1e64345763b0134c5d

Observation 65229ffc-8d53-46e3-b205-549ceff1b21c · outbound

This paper cites In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion

Reference 8

Resolution
verified exact
doi, observed 2026-08-16T11:53:29.036959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.687722Z digest=sha256:f0f7da2a25df107042335b6e4ec03fefb926f709b731b87243f1a4b382128207

Observation 8b9976a4-4f52-4bd9-aa9d-cf85a28bd283 · outbound

This paper cites Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.692181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.692181Z digest=sha256:22dd74f6f98963189f9c434e2f1faa7bc747b844ffad77cc2e2e977e67e92156

Observation f00bda77-aa71-440a-8c60-b776ea5265cf · outbound

This paper cites arXiv preprint arXiv:2407.06153 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.06153 (2024)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.696732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.696732Z digest=sha256:09f5716c4f9ed00426708656517392c638d0438998074e3d5b2cb4cb2f2c9d7a

Observation ae96fd29-0008-44bd-beab-b6cc578ff7bf · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.299883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.729125Z digest=sha256:0c1fd7657e86f5b1e4329f561e500ee5fd23151d1ea207bdb69477e1933d5f3b

Observation ce400777-20f3-462d-ac3f-b7f3fcd0fb7a · outbound

This paper cites MetaRM: Shifted Distributions Alignment via Meta-Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay MetaRM: Shifted Distributions Alignment via Meta-Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.844350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.844350Z digest=sha256:ed2f78e85047aaf10ecd30bee6978e005adbad6c4fefd47ff6ea31d4223e7194

Observation 03c2a41e-5bc4-4e4c-a7df-49adf5076bd3 · outbound

This paper cites Towards Understanding the Capability of Large Language Models on Code Clone Detection: A Survey.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Understanding the Capability of Large Language Models on Code Clone Detection: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.946663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.946663Z digest=sha256:cc749df275ac62c405bb59038b3d12daa4d6acc0162c6612a4cb06057a621d0d

Observation c00eeada-e0ed-49f8-86fe-a6ec1ce69bcc · outbound

This paper cites arXiv preprint arXiv:2506.02672 (2025).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2506.02672 (2025)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.950609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.950609Z digest=sha256:8c5341adda7fb9215e869551138f703e47582bb199451b785a99c46332fe6294

Observation 3ded2841-43b0-4cf3-9a7e-c520a900f59e · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the AssociationforComputationalLinguistics(Volume1:LongPapers).pp.1932–1945 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the AssociationforComputationalLinguistics(Volume1:LongPapers).pp.1932–1945 (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.288728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.954831Z digest=sha256:1e7d75831163da757f43d50d3037cd19a4fcb236b941128a90e9ee4c3b1b5307

Observation a4f13fcd-9947-40fd-92d0-73e41fcf66b6 · outbound

This paper cites Nature590, 580 – 586 (2020),https://api.semanticscholar.org/ CorpusID:216552951.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Nature590, 580 – 586 (2020),https://api.semanticscholar.org/ CorpusID:216552951

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.276945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.957910Z digest=sha256:8250719e38c1cef49475e637acb555dc8aa8268adffc61b6cb2c593a4d229ceb

Observation 265801ab-340f-4118-a497-9cccdf90e407 · outbound

This paper cites In: Proceedings of the 41st International Conference on Machine Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 41st International Conference on Machine Learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.231147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.961503Z digest=sha256:bc3160dec23956362690b3880b8836e94d99c7ffa88cde1b1ce93c1e3b1fd1f6

Observation b5a32ec2-5821-47b7-bbf0-a158e84ef373 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.965567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.965567Z digest=sha256:91e577d596c6ff8a34283df896f5097218fb28fb74959fc7f47eeaf3d75aa8e2

Observation eefee861-05aa-4a28-b84a-77bf10127813 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:30.130543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.969974Z digest=sha256:e17515fb54d4cff3e9e76c1f05d3c05f2c17b6804722994b7e346eeb940dc7d5

Observation 48a7c01a-d34b-4642-a506-bddcf65291cd · outbound

This paper cites In: Vanschoren, J., Yeung, S.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Vanschoren, J., Yeung, S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.118878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:27.973774Z digest=sha256:50c4cd5821818bef946eb9f3763f38e78c76bc24b3b044f0cdeac0bdc5c904cb

Observation 7b9dcbfb-6158-4163-9e86-7b9800b51638 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.977900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.977900Z digest=sha256:eea96d8af7b5729fbe09b0547ee22a5abf90d11ed91ba488cbc327ccf79fdc6c

Observation 9055ac67-6aae-419c-95a0-8be5bed2262c · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Reasoning in Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.013293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.013293Z digest=sha256:abd1c84f885292dde1b025c163071ea5f89e0da8296b5a256a2076685a074c39

Observation 7a2d2248-5edb-4f01-a135-7b470cc0818d · outbound

This paper cites RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:53:29.552762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.109352Z digest=sha256:9adb5c6a2fb195d174806600a2ee0c14be9ab15133ef0eccf61b8729ed92f6ae

Observation 2e2a4f09-b6db-4598-92c3-e9807baa9d1a · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.162206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.162206Z digest=sha256:7156c503cac63a4571968441e0d5a9366c278a7353011305b8ac2153301eaf13

Observation 25a5fd15-d033-42ca-bac4-4b1d4b32e08d · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.167820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.167820Z digest=sha256:45ea631782abca610533a6690d200b694b83caa52b56077e15b7551d2e565cc5

Observation b9c4d114-ae70-4143-b99d-e46c18016fa9 · outbound

This paper cites Information Fusion85, 1–22 (2022).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Information Fusion85, 1–22 (2022)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.174446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.174446Z digest=sha256:9b32462ce1eeed5f61c5a699b43c7364d338800a1dcdbbf4d4c4d729a06defc0

Observation 1896f1dd-c6fb-48f3-b105-f6967430cbcd · outbound

This paper cites StarCoder: may the source be with you!.

Improving RL Exploration for LLM Reasoning through Retrospective Replay StarCoder: may the source be with you!

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.180369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.180369Z digest=sha256:0ae5ef4d2d303e90769a1e8c37484b9ee9e09f9eef0cf1fc123d3f42d1dca3e4

Observation f60984d5-12cf-40a0-889c-d2d0327c1bd5 · outbound

This paper cites Deep Reinforcement Learning: An Overview.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Deep Reinforcement Learning: An Overview

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.185715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.185715Z digest=sha256:f230c58fd9e9bf08789284f107abdddf520f0752278073307722bb3b2846c008

Observation d67fe6a5-165f-4abe-981a-cdc3bc643bc2 · outbound

This paper cites RLTF: Reinforcement Learning from Unit Test Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay RLTF: Reinforcement Learning from Unit Test Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.190874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.190874Z digest=sha256:49149b21aeca95e0bb7297a352ce718d9b8a9d558dcf984a51f7b716b99967e3

Observation b834e9ad-e8c2-4d10-8381-da9a0ba90d6f · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Improving RL Exploration for LLM Reasoning through Retrospective Replay WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.195335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.195335Z digest=sha256:b48f2853a642a504322cb81b880a97e48d981b37c72621b58c9f48e6f0f0a351

Observation 4196ed8a-44ca-45b0-8c1b-b0e3c47716de · outbound

This paper cites Brief analysis of DeepSeek R1 and its implications for Generative AI.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Brief analysis of DeepSeek R1 and its implications for Generative AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.200209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.200209Z digest=sha256:87852be59a7450a43c8121211104725dbbed4ce7fdb79f572d8f308f4387b5e9

Observation ec987063-8bc8-4f1c-8f08-1107e78c04e4 · outbound

This paper cites In: 2018 IEEE Inter- national Conference on Robotics and Automation (ICRA).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2018 IEEE Inter- national Conference on Robotics and Automation (ICRA)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.258352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.258352Z digest=sha256:195b29e9e8ed0bc4c3da0b9d0bc157f840e4fc3e65303be3ea61708685137c66

Observation b74eb59b-54e3-4ada-ae5b-0b7b012fce14 · outbound

This paper cites arXiv preprint arXiv:2407.11511 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.11511 (2024)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.394078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.394078Z digest=sha256:217c3c85701e8957e6c61905f370ff96edc92dd34ad6c148d145847893b40b6b

Observation de719de5-add6-4a7a-aa2d-16575cc7cd27 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Code Llama: Open Foundation Models for Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.457662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.457662Z digest=sha256:fdf25e979ed1c689da4dd92c100f1864d222c64d95c5fb9f47f878ee09e71099

Observation 98c32d55-ab36-466a-a87f-bc3f7c533e1b · outbound

This paper cites Prioritized Experience Replay.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Prioritized Experience Replay

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.462546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.462546Z digest=sha256:b01b878237f4029fc6df7f95749a163ec78bd4f50539540495dde32f2d1fec5d

Observation 667eed7b-7bc1-4c79-9e4e-437d47d7f5e1 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Improving RL Exploration for LLM Reasoning through Retrospective Replay High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.512230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.512230Z digest=sha256:e255b895e4aab7c5e4c21dcb84cb9a0e5aa32b20ca2825e92bc42d09796920f9

Observation 52cc8e12-af0e-4e6a-b088-e953c55ea7b2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.596337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.596337Z digest=sha256:6b1137d1ddbd04b2ac0e11592c39f636cdf68fe5911376eff75469048715ea16

Observation 9db5262c-478d-4508-abde-356ab74bef36 · outbound

This paper cites 5-thinking: Advancing superb reasoning models with reinforcement learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay 5-thinking: Advancing superb reasoning models with reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.600249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.600249Z digest=sha256:1d97736fafab3cdb0181b095eaf12dbce8d583724f3a84974cf82367a71abf37

Observation f4bc0226-3566-42ed-aba4-52742dfa2fda · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.603999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.603999Z digest=sha256:097f675ddde3b7b76465eec8ae72f523c3888d63f7ddf2dfa7529ff0bacf35a5

Observation f232fc04-9978-43c8-8f7e-0226eb9bcb16 · outbound

This paper cites In: The 2023 Conference on Empirical Methods in Natural Language Processing.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: The 2023 Conference on Empirical Methods in Natural Language Processing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.101836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.608387Z digest=sha256:6551d9d17ec397a5df69c012569a43fb5577574153f60b33ea971f2981dbd445

Observation dfe5a085-9bf6-49bb-bcf9-4e0876a6106a · outbound

This paper cites Execution-based Code Generation using Deep Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Execution-based Code Generation using Deep Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.614182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.614182Z digest=sha256:16bfeb7d39952d4da0d944a7872594a5641404bbab1f4709b157705b913ae2c1

Observation 85e0fc10-7300-46e5-98b4-581cc48cf3a4 · outbound

This paper cites Advances in neural informa- tion processing systems12 (1999).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Advances in neural informa- tion processing systems12 (1999)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.047980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.689683Z digest=sha256:ea0fce42e3a3405bbb1f86f0254c7871d5f7bb6b50b79b9c922213d0c7e7ea65

Observation 562cef5c-f270-45ac-8046-207313ef9b73 · outbound

This paper cites Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.697223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.697223Z digest=sha256:a60e2509cddaa20fdccc713cfe7ebf4dbfe7dce620053b8d1dc1ca153f97108c

Observation 32e0b9d7-d87f-4c01-975a-4981ca8b7e02 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:29.901092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.702216Z digest=sha256:9f2f5e4da220b30dc2eaf6aaa80c1b063a646a3826446c32c1b9abaf576309af

Observation 32f7067e-076a-4510-b437-f9e204aa0d67 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.707518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.707518Z digest=sha256:679da5c66160820e885679e998cf6e92fbc40cf5e5d42b4ff0bb91bc63fb6826

Observation 3a8e414d-0b82-4ddc-a9a5-3eaa050741c2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Solving math word problems with process- and outcome-based feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.711624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.711624Z digest=sha256:67387afb1805ff109cefd56312a704ea7bde38eb09e9b4d83e10cb47222ad37e

Observation 6a10ab00-4e0a-4284-89eb-e8ac65cf2a98 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.737965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.737965Z digest=sha256:889871f8a1cbac8fcb80f868834ebb0210c09b0b4247956296ebed211eb973e5

Observation 61e3575d-57e1-43c4-ac9a-d246f320caf0 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.847092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.847092Z digest=sha256:6a31de4d6c1bc9cef720e53f08ccf6a019e1b40a9db415797dc75bbfc5017f51

Observation 28577e6f-3437-4add-90e9-300765c13701 · outbound

This paper cites Machine learning8, 229–256 (1992).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Machine learning8, 229–256 (1992)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.851183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.851183Z digest=sha256:1f392b80b9c2545e2e98b51a0c8291126b3fd45bf715255f295c5508bbb956a6

Observation c4fcf402-90c7-4a45-a758-3bc594580228 · outbound

This paper cites Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.855634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.855634Z digest=sha256:0daa77328adf117dbb4618949c78f244412e9f9a4d67a97f017b48f67470d921

Observation a571a8ea-e785-4453-ad81-2414c8c6a3d3 · outbound

This paper cites In: International Conference on Machine Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: International Conference on Machine Learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:29.883416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.860151Z digest=sha256:a7156722926bf63c96f446f06c2cc150d2fc68f04a541152d3db0d111fe2d054

Observation a016837c-4a05-42e1-b30e-e793c6fd24c5 · outbound

This paper cites Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.864222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.864222Z digest=sha256:9ed63c15cd0df41857f5a0c4b73547c99db18c73b77df23c7a94d7782cf7abcc

Observation ede8af13-74e9-4136-b940-c3d979966938 · outbound

This paper cites arXiv preprint arXiv:2505.17793 (2025).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2505.17793 (2025)

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.869017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.869017Z digest=sha256:9dff920dac3bc5701fc84a3292e2f201dc1cc57a60bd7a3629c27bb79859acce

Observation 455a8c2b-08a9-48e3-9529-729c736bc6d6 · outbound

This paper cites A Deeper Look at Experience Replay.

Improving RL Exploration for LLM Reasoning through Retrospective Replay A Deeper Look at Experience Replay

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.872617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.872617Z digest=sha256:57d3f2d3f86248b53bc6ab854d451f622dd9d15251929c41635cf6555ff4f8ea

Observation bff13c8d-45e2-4636-a8c6-262d0a021569 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:29.872604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:28.928556Z digest=sha256:5238daccc5fdd07afee61fe07f5a3aa15f2dfacdbb7b92ea6e8d9663a171264c

Observation 6421c666-6154-4714-ba74-d565060ac54b · outbound

This paper cites In: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (2023).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (2023)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:29.859319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:53:29.001062Z digest=sha256:67296cad9def38ed7e225f7dc087fd8778cbf8e76776f1f8afcf77888700636e

Pith citing papers

Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · inbound

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning cites this paper.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.055667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.055667Z digest=sha256:a23e857f6860e65b05760b3860df5ffb707262c9e44ab9cb07c6e1ce701e55f6

Observation 331e8f51-bedd-44b1-8469-cd6695202c97 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.295695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:7ae1cb3bcaac34fece8b07fd52db6e581597030cad9676a142b150e02d7045e7

Observation 8d570112-a7fc-4c0c-84ae-491828bceca0 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:12.821899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:1844fa056a38f4467c5abef42c2f26e597610b4bd5817c7a41c422c85ad983eb

Observation 0d92ba2d-2def-4264-a3d6-9daafaa5fb0a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 266

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.562994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:27233b1ec734eeeab2a0d53f07d6080f555bb35cf7c8dc6bc3001f0c8f1dffdd

Observation f71fc0eb-23d5-43eb-abf0-9f60c6da2b44 · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.167562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:a2d7255348a0ad6821acd77a6d5b9d4a72ccaadbb137f2b2f78bd5e8188411f5

Observation f3695488-aca2-48e0-b879-1b07bbbd95a5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.190471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:06fdf2674b0bed5e3a571746358c5ec84e5fdca01a3f0455704cfdfdc16b4308