Pith. sign in

Paper Citation Record · LEDGER

Improving RL Exploration for LLM Reasoning through Retrospective Replay

As of 19 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 6 inbound Pith citation observations for arXiv:2504.14363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14363 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:29.001062Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.055667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.189087Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 888524d1-e700-4d13-b272-6e7915a7e389 · outbound

This paper cites IEEE Signal Processing Magazine 34(6), 26–38 (2017).

Improving RL Exploration for LLM Reasoning through Retrospective Replay IEEE Signal Processing Magazine 34(6), 26–38 (2017)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.505944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.505944Z digest=sha256:a1bd686b9a93129c57c50e6b42fe61d3ad1abdd33a5dd7719badcb4a2b597987

Observation dae19bd0-8887-42d9-9b17-6e1071ebc1e6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.575439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.575439Z digest=sha256:ac9a617a84f72e8436be329a4c9583a2892046fb317e52a77f55793b94d7c4d2

Observation c90a7971-3e79-4e58-bc1b-850676693ab8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.666657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.666657Z digest=sha256:a6cd60cf9a8420a56a5900ec91506e0045a8492699176c5d19f8dc10989e4662

Observation 9c39bf08-9c78-4d61-9ebe-bcffa6d48bb5 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.671172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.671172Z digest=sha256:6369f3650c8e0ea2d282b817b0a76146cfca96917afd46ed3babb0b22b58d41a

Observation d20b8933-9a52-46b1-bfc6-e781c9329f0a · outbound

This paper cites DocFusion: A Unified Framework for Document Parsing Tasks.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DocFusion: A Unified Framework for Document Parsing Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.675801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.675801Z digest=sha256:7cc0921bb01e6ffae47b2b4bc9ed6cd373171863d6a9fe0e81ab8f061dcc13e0

Observation 42bc350b-17e5-423b-92a7-b384c0860eac · outbound

This paper cites PanGu-Coder: Program Synthesis with Function-Level Language Modeling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay PanGu-Coder: Program Synthesis with Function-Level Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.679663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.679663Z digest=sha256:374e507e958d94734a41615eac5562cd89fab1c22252f042701267796909e3ac

Observation ad4659b5-0974-46d4-a975-452382807278 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.684280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.684280Z digest=sha256:dff75dff0e25018a8ac04e142a1584b9b45317a9e5975f1e64345763b0134c5d

Observation 65229ffc-8d53-46e3-b205-549ceff1b21c · outbound

This paper cites In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion

Reference 8

Resolution
verified exact
doi, observed 2026-08-16T11:53:29.036959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.687722Z digest=sha256:19ae77cb9245156390c058791fd938bbf7005d21db214bd01755a9ce4dfd5c69

Observation 8b9976a4-4f52-4bd9-aa9d-cf85a28bd283 · outbound

This paper cites Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.692181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.692181Z digest=sha256:22dd74f6f98963189f9c434e2f1faa7bc747b844ffad77cc2e2e977e67e92156

Observation f00bda77-aa71-440a-8c60-b776ea5265cf · outbound

This paper cites arXiv preprint arXiv:2407.06153 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.06153 (2024)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.696732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.696732Z digest=sha256:09f5716c4f9ed00426708656517392c638d0438998074e3d5b2cb4cb2f2c9d7a

Observation ae96fd29-0008-44bd-beab-b6cc578ff7bf · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.299883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.729125Z digest=sha256:b0738cf461b9c6d070a8b692028df6410c737ecc9d495160f58ca9b477b70bef

Observation ce400777-20f3-462d-ac3f-b7f3fcd0fb7a · outbound

This paper cites MetaRM: Shifted Distributions Alignment via Meta-Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay MetaRM: Shifted Distributions Alignment via Meta-Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.844350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.844350Z digest=sha256:c373a16aa79dd2950ad3b82b5717f773d399aaea69a7e8012a280a3f7faac345

Observation 03c2a41e-5bc4-4e4c-a7df-49adf5076bd3 · outbound

This paper cites Towards Understanding the Capability of Large Language Models on Code Clone Detection: A Survey.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Understanding the Capability of Large Language Models on Code Clone Detection: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.946663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.946663Z digest=sha256:cc749df275ac62c405bb59038b3d12daa4d6acc0162c6612a4cb06057a621d0d

Observation c00eeada-e0ed-49f8-86fe-a6ec1ce69bcc · outbound

This paper cites arXiv preprint arXiv:2506.02672 (2025).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2506.02672 (2025)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.950609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.950609Z digest=sha256:8c5341adda7fb9215e869551138f703e47582bb199451b785a99c46332fe6294

Observation 3ded2841-43b0-4cf3-9a7e-c520a900f59e · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the AssociationforComputationalLinguistics(Volume1:LongPapers).pp.1932–1945 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the AssociationforComputationalLinguistics(Volume1:LongPapers).pp.1932–1945 (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.288728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.954831Z digest=sha256:79fefc2f97e3b44efb9714a01692be7304aa19c2591658fc05565279b7c5ba70

Observation a4f13fcd-9947-40fd-92d0-73e41fcf66b6 · outbound

This paper cites Nature590, 580 – 586 (2020),https://api.semanticscholar.org/ CorpusID:216552951.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Nature590, 580 – 586 (2020),https://api.semanticscholar.org/ CorpusID:216552951

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.276945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.957910Z digest=sha256:bc39b24b5a272cc6386e2a829e262287f0bb6711c1d5714d68014502a022cd78

Observation 265801ab-340f-4118-a497-9cccdf90e407 · outbound

This paper cites In: Proceedings of the 41st International Conference on Machine Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 41st International Conference on Machine Learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.231147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.961503Z digest=sha256:fc0f8bb49f4e86fced8278c6fb37df8be06d108e609562c54c2af205215b419f

Observation b5a32ec2-5821-47b7-bbf0-a158e84ef373 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.965567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.965567Z digest=sha256:91e577d596c6ff8a34283df896f5097218fb28fb74959fc7f47eeaf3d75aa8e2

Observation eefee861-05aa-4a28-b84a-77bf10127813 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:30.130543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.969974Z digest=sha256:f3014ebf41f6ddd4a0d7437a68983c2b758cff33f89920d8bcffb8e4ddf1bc31

Observation 48a7c01a-d34b-4642-a506-bddcf65291cd · outbound

This paper cites In: Vanschoren, J., Yeung, S.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Vanschoren, J., Yeung, S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.118878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:27.973774Z digest=sha256:bd313ab99ed45a051175e6fa18cf8a3a4364cc255a2e22572dad3efb5438f097

Observation 7b9dcbfb-6158-4163-9e86-7b9800b51638 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.977900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.977900Z digest=sha256:eea96d8af7b5729fbe09b0547ee22a5abf90d11ed91ba488cbc327ccf79fdc6c

Observation 9055ac67-6aae-419c-95a0-8be5bed2262c · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Reasoning in Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.013293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.013293Z digest=sha256:abd1c84f885292dde1b025c163071ea5f89e0da8296b5a256a2076685a074c39

Observation 7a2d2248-5edb-4f01-a135-7b470cc0818d · outbound

This paper cites RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:53:29.552762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.109352Z digest=sha256:7fb9a388ebf84f749ad8581beb9a9c80902d8f2baed665d1a3eea2a34fbe8e0f

Observation 2e2a4f09-b6db-4598-92c3-e9807baa9d1a · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.162206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.162206Z digest=sha256:7156c503cac63a4571968441e0d5a9366c278a7353011305b8ac2153301eaf13

Observation 25a5fd15-d033-42ca-bac4-4b1d4b32e08d · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.167820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.167820Z digest=sha256:45ea631782abca610533a6690d200b694b83caa52b56077e15b7551d2e565cc5

Observation b9c4d114-ae70-4143-b99d-e46c18016fa9 · outbound

This paper cites Information Fusion85, 1–22 (2022).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Information Fusion85, 1–22 (2022)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.174446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.174446Z digest=sha256:9b32462ce1eeed5f61c5a699b43c7364d338800a1dcdbbf4d4c4d729a06defc0

Observation 1896f1dd-c6fb-48f3-b105-f6967430cbcd · outbound

This paper cites StarCoder: may the source be with you!.

Improving RL Exploration for LLM Reasoning through Retrospective Replay StarCoder: may the source be with you!

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.180369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.180369Z digest=sha256:0ae5ef4d2d303e90769a1e8c37484b9ee9e09f9eef0cf1fc123d3f42d1dca3e4

Observation f60984d5-12cf-40a0-889c-d2d0327c1bd5 · outbound

This paper cites Deep Reinforcement Learning: An Overview.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Deep Reinforcement Learning: An Overview

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.185715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.185715Z digest=sha256:f230c58fd9e9bf08789284f107abdddf520f0752278073307722bb3b2846c008

Observation d67fe6a5-165f-4abe-981a-cdc3bc643bc2 · outbound

This paper cites RLTF: Reinforcement Learning from Unit Test Feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay RLTF: Reinforcement Learning from Unit Test Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.190874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.190874Z digest=sha256:49149b21aeca95e0bb7297a352ce718d9b8a9d558dcf984a51f7b716b99967e3

Observation b834e9ad-e8c2-4d10-8381-da9a0ba90d6f · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Improving RL Exploration for LLM Reasoning through Retrospective Replay WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.195335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.195335Z digest=sha256:b48f2853a642a504322cb81b880a97e48d981b37c72621b58c9f48e6f0f0a351

Observation 4196ed8a-44ca-45b0-8c1b-b0e3c47716de · outbound

This paper cites Brief analysis of DeepSeek R1 and its implications for Generative AI.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Brief analysis of DeepSeek R1 and its implications for Generative AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.200209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.200209Z digest=sha256:87852be59a7450a43c8121211104725dbbed4ce7fdb79f572d8f308f4387b5e9

Observation ec987063-8bc8-4f1c-8f08-1107e78c04e4 · outbound

This paper cites In: 2018 IEEE Inter- national Conference on Robotics and Automation (ICRA).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2018 IEEE Inter- national Conference on Robotics and Automation (ICRA)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.258352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.258352Z digest=sha256:195b29e9e8ed0bc4c3da0b9d0bc157f840e4fc3e65303be3ea61708685137c66

Observation b74eb59b-54e3-4ada-ae5b-0b7b012fce14 · outbound

This paper cites arXiv preprint arXiv:2407.11511 (2024).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.11511 (2024)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.394078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.394078Z digest=sha256:217c3c85701e8957e6c61905f370ff96edc92dd34ad6c148d145847893b40b6b

Observation de719de5-add6-4a7a-aa2d-16575cc7cd27 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Code Llama: Open Foundation Models for Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.457662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.457662Z digest=sha256:b59b1c0cb7ab930ccf09e9c63c1c74aab6baefce6e23ff31e64de6c0ebcdbd22

Observation 98c32d55-ab36-466a-a87f-bc3f7c533e1b · outbound

This paper cites Prioritized Experience Replay.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Prioritized Experience Replay

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.462546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.462546Z digest=sha256:b01b878237f4029fc6df7f95749a163ec78bd4f50539540495dde32f2d1fec5d

Observation 667eed7b-7bc1-4c79-9e4e-437d47d7f5e1 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Improving RL Exploration for LLM Reasoning through Retrospective Replay High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.512230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.512230Z digest=sha256:e255b895e4aab7c5e4c21dcb84cb9a0e5aa32b20ca2825e92bc42d09796920f9

Observation 52cc8e12-af0e-4e6a-b088-e953c55ea7b2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.596337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.596337Z digest=sha256:6b1137d1ddbd04b2ac0e11592c39f636cdf68fe5911376eff75469048715ea16

Observation 9db5262c-478d-4508-abde-356ab74bef36 · outbound

This paper cites 5-thinking: Advancing superb reasoning models with reinforcement learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay 5-thinking: Advancing superb reasoning models with reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.600249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.600249Z digest=sha256:1d97736fafab3cdb0181b095eaf12dbce8d583724f3a84974cf82367a71abf37

Observation f4bc0226-3566-42ed-aba4-52742dfa2fda · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.603999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.603999Z digest=sha256:097f675ddde3b7b76465eec8ae72f523c3888d63f7ddf2dfa7529ff0bacf35a5

Observation f232fc04-9978-43c8-8f7e-0226eb9bcb16 · outbound

This paper cites In: The 2023 Conference on Empirical Methods in Natural Language Processing.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: The 2023 Conference on Empirical Methods in Natural Language Processing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.101836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.608387Z digest=sha256:f7e3d0e0120b658189de7afd20a20f2dc1a6a3c72ddec5f4d5862e62ccb3c6b5

Observation dfe5a085-9bf6-49bb-bcf9-4e0876a6106a · outbound

This paper cites Execution-based Code Generation using Deep Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Execution-based Code Generation using Deep Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.614182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.614182Z digest=sha256:16bfeb7d39952d4da0d944a7872594a5641404bbab1f4709b157705b913ae2c1

Observation 85e0fc10-7300-46e5-98b4-581cc48cf3a4 · outbound

This paper cites Advances in neural informa- tion processing systems12 (1999).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Advances in neural informa- tion processing systems12 (1999)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:30.047980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.689683Z digest=sha256:78bd90e56be8bc349e14528710f078e14b0e96a72807d96bfe033b25b8e85f12

Observation 562cef5c-f270-45ac-8046-207313ef9b73 · outbound

This paper cites Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.697223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.697223Z digest=sha256:a60e2509cddaa20fdccc713cfe7ebf4dbfe7dce620053b8d1dc1ca153f97108c

Observation 32e0b9d7-d87f-4c01-975a-4981ca8b7e02 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:29.901092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.702216Z digest=sha256:49162157bfd9519999a512216eb6e6d74c155cb3d84f1c1f4fbffbc67379262d

Observation 32f7067e-076a-4510-b437-f9e204aa0d67 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.707518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.707518Z digest=sha256:679da5c66160820e885679e998cf6e92fbc40cf5e5d42b4ff0bb91bc63fb6826

Observation 3a8e414d-0b82-4ddc-a9a5-3eaa050741c2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Solving math word problems with process- and outcome-based feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.711624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.711624Z digest=sha256:67387afb1805ff109cefd56312a704ea7bde38eb09e9b4d83e10cb47222ad37e

Observation 6a10ab00-4e0a-4284-89eb-e8ac65cf2a98 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.737965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.737965Z digest=sha256:889871f8a1cbac8fcb80f868834ebb0210c09b0b4247956296ebed211eb973e5

Observation 61e3575d-57e1-43c4-ac9a-d246f320caf0 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.847092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.847092Z digest=sha256:0cf44aeba4a85da946e5bad15e9aaf0ee14d059570d30066af6aef43e37dddfd

Observation 28577e6f-3437-4add-90e9-300765c13701 · outbound

This paper cites Machine learning8, 229–256 (1992).

Improving RL Exploration for LLM Reasoning through Retrospective Replay Machine learning8, 229–256 (1992)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.851183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.851183Z digest=sha256:1f392b80b9c2545e2e98b51a0c8291126b3fd45bf715255f295c5508bbb956a6

Observation c4fcf402-90c7-4a45-a758-3bc594580228 · outbound

This paper cites Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.855634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.855634Z digest=sha256:0daa77328adf117dbb4618949c78f244412e9f9a4d67a97f017b48f67470d921

Observation a571a8ea-e785-4453-ad81-2414c8c6a3d3 · outbound

This paper cites In: International Conference on Machine Learning.

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: International Conference on Machine Learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:29.883416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.860151Z digest=sha256:34f1bfbd8f5099d782eddcd782d3a95cca0be25817ac2a46352bcfee2e7428c8

Observation a016837c-4a05-42e1-b30e-e793c6fd24c5 · outbound

This paper cites Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.864222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.864222Z digest=sha256:9ed63c15cd0df41857f5a0c4b73547c99db18c73b77df23c7a94d7782cf7abcc

Observation ede8af13-74e9-4136-b940-c3d979966938 · outbound

This paper cites arXiv preprint arXiv:2505.17793 (2025).

Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2505.17793 (2025)

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.869017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.869017Z digest=sha256:9dff920dac3bc5701fc84a3292e2f201dc1cc57a60bd7a3629c27bb79859acce

Observation 455a8c2b-08a9-48e3-9529-729c736bc6d6 · outbound

This paper cites A Deeper Look at Experience Replay.

Improving RL Exploration for LLM Reasoning through Retrospective Replay A Deeper Look at Experience Replay

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:28.872617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:28.872617Z digest=sha256:57d3f2d3f86248b53bc6ab854d451f622dd9d15251929c41635cf6555ff4f8ea

Observation bff13c8d-45e2-4636-a8c6-262d0a021569 · outbound

This paper cites an unresolved cited work.

Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:53:29.872604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:28.928556Z digest=sha256:9bc167af669084fd95f4f70eda27506ed93502636b255e02ba90d84f2929e900

Observation 6421c666-6154-4714-ba74-d565060ac54b · outbound

This paper cites In: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (2023).

Improving RL Exploration for LLM Reasoning through Retrospective Replay In: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (2023)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:29.859319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:53:29.001062Z digest=sha256:70383549e0281bbc8adaf02d3c64f25de2ab98c5b53c1b309f0246e04766cbe7

Pith citing papers

Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · inbound

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning cites this paper.

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:38.055667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:38.055667Z digest=sha256:a23e857f6860e65b05760b3860df5ffb707262c9e44ab9cb07c6e1ce701e55f6

Observation 331e8f51-bedd-44b1-8469-cd6695202c97 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.295695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:68e3b14d2618ee9c452efb513bf75aba5a259981649f6630d40a2c4d1c8c92e1

Observation 8d570112-a7fc-4c0c-84ae-491828bceca0 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:12.821899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:6d8e314e8ad07c0eb28c7aebd121ddbc9c1e0365f9a92d580133b618b162069b

Observation 0d92ba2d-2def-4264-a3d6-9daafaa5fb0a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 266

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.562994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:bb6cee9b296fb0d815971206597c431ba42c04fd150d224af03b4624ba558bc3

Observation f71fc0eb-23d5-43eb-abf0-9f60c6da2b44 · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.167562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:edf8526387381444f82890ecc26e7e17c530640e57ce5a793365fe4d5cae3737

Observation f3695488-aca2-48e0-b879-1b07bbbd95a5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.190471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:e301b96bd35c7f0bc8e7848c95b0dc19ef19fffa88106ee89fd978f9401f1f42