Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:34:31.715954Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2604.11297.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:34:31.715954Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ea33fd6-dc3d-403b-ba4f-6e4cd7ea5b54 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dde50c55-7418-4b3a-a646-fa3c4f5f4ab2 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b21c6ace-e1f4-4553-8d50-6b21cb945071 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2d69158-f2b0-4ad5-94c3-01f4f57f2a5b · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Execution-basedcodegenerationusingdeep reinforcement learning.Trans
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d640a3bd-07c7-4972-badd-a7b0503d3907 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Christiano, Jan Leike, and Ryan Lowe
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46665d13-9645-4259-a5b0-37c611dc06d0 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping arXiv preprint arXiv:2503.01067 , year=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dd6fcf5-e8da-4ec2-9e1e-f8ca91e0f7d9 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7aa33035-2676-4431-b5d7-fe5e91bf9747 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Expected return causes outcome-level mode collapse in re- inforcement learning and how to fix it with inverse probability scaling.CoRR, abs/2601.21669
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d7b1b52-dcec-47ca-a772-3eab1aab6467 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b09873b-c376-45a9-a2e5-0add5efe78e1 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92cfaa51-ace3-4881-8f16-8b0abdee0d69 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1796ad6a-0103-4283-90e8-518143746bb9 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DSDR: Dual-scale diversity regularization for exploration in LLM reasoning.arXiv preprint arXiv:2602.19895
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50586630-192a-4d37-b726-4078121ca7b1 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e15d907-0407-4af1-8e92-1c9507d29ac0 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The neural basis of human error processing: Reinforcement learning, dopamine, and the error-related negativity.Psychological Review, 109:679–709, 11 2002
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b6c280-e1b0-4950-9013-890dcc8272ee · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The Journal of Open Source Software 2(11) (mar 2017)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72752f53-1af9-4691-bfa3-3f3127f297b4 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ccf4e04-f567-40fb-898c-0cad0c1e04c7 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Qimeng-codev-r1: Reasoning-enhanced verilog generation.arXiv preprint arXiv:2505.24183
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70c70560-a7fa-4a35-906f-cfabf5963d6e · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2910a4f1-355f-4a94-8ebf-9d767a333652 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea46ed91-eb05-49a5-a9f7-660f2acd8fd1 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping REARANK: reasoning re-ranking agent via reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3995b01b-fa87-4716-b40c-8a51fc70c80a · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Let’s verify step by step
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29894fa3-4f98-402c-bbf6-9272ddeb74cf · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Solving math word problems with process- and outcome-based feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c326e03-d676-43c8-b538-85efe56b9e3f · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c099b7f6-9a2e-453e-adf0-ac095ad01702 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Rewarding the rare: Uniqueness-aware rl for creative problem solving in llms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f61e66d-647e-4ac4-9bb8-3c2b623fc48a · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec353b52-5182-4790-b97c-e9927c819038 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 176ffacb-3cee-4ccf-8105-48c887858eba · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Emogen: Emotional image content generation with text-to-image diffusion models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80a1572b-eebf-4928-8bc6-9f159a55171c · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Multi-objective evolution of heuristic usinglargelanguagemodel
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a949e1ae-6f48-48d1-9a7e-922d7fcce932 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Latent reward: Llm-empowered credit assignment in episodic reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0beafab-6a79-4067-87b4-80c40e6eec9d · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Revolve: Reward evolution with large language models using human feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31af2722-e6e6-499a-808b-eaf668f1bd25 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Eureka: Human-level reward design via coding large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ae3172b-5fe5-41e2-93fd-551c5931cbc1 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebfd1f85-9597-4a85-ac28-acc07ae2d4b6 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Locating and editing factual associations in GPT
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b915ce-fee4-4436-8fb4-f90d7f577115 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Daniel Freeman, Theodore R
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b0d8499-8447-4727-8b51-1ba97563fe5e · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping In-context learning and induction heads.Transformer Circuits Thread
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c39b9499-16e6-4950-b3df-ba9b0e5c3b20 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01801f0c-9b33-4699-b891-51ecca4002d3 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c865a2e-4829-4c33-9139-64133ac26ff9 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Verifyingchain-of-thought reasoning via its computational graph.CoRR, abs/2510.09312
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74c4eec3-f643-4301-b029-6ecfe88d43e0 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f4da862-42a4-4883-92f6-6af5516f5b65 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 519e5c6a-547b-473c-88ef-f169412f3fb6 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Measuring mathematical problem solving with the MATH dataset
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17af816a-4cc0-47c4-ba6a-27210a2467d3 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Understanding R1-Zero-Like Training: A Critical Perspective
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f26a4de9-4d78-4f9b-809c-ff1a8440c331 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Hybridflow: A flexible and efficient RLHF framework
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcac0cb1-e92d-4dbc-aed3-164c6c91d074 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Hybridflow: A flexible and efficient rlhf framework
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de4032af-9a3c-410e-b58e-aabe62879214 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Reasoning with Exploration: An Entropy Perspective
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f09b3a45-9878-4089-908f-0326eac4c8e8 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13(9):9
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9515d54-e8fe-48c1-9bbd-5606c72efdf5 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur- Ari, and Vedant Misra
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d979fce-46d0-43f8-8d39-f56e3b96a242 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4e071d3-2b33-455d-ae8e-8d1518f2d3f1 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping 1 𝐾 𝐾Õ 𝑖=1 min 𝑟𝑖(𝜃)𝐴𝑖 ,clip(𝑟𝑖(𝜃),1−𝜖,1+𝜖)𝐴𝑖 !# −𝛽𝔻KL[𝜋𝜃∥𝜋ref] (1) ℒDAPO(𝜃)=𝔼𝑞∼𝒟
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1a15149-0883-499e-970d-7a5a053e60ee · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Since89 2 is closer to 8085, we use it: √ 8085≈89.9166, and thus: 𝑝= −1+89.9166 2 ≈88.9166 2 ≈44.4583, which isn’t an integer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9f9fc8c-6c00-424e-8b7e-46cdb7af9e72 · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The factors of𝑝2 are1, 𝑝, and𝑝 2
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad69b496-dc57-4420-9395-118cc600c70b · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping "" Returns a sorted list of all divisors of n
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c47dc01-abe8-490c-90d7-7053fcedcbcb · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The sum of these three divisors is1+𝑑+𝑛 𝑑 =2022
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b250739-dda4-4d81-8bfe-5265199745fd · outbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping reasoning path
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.