Pith. sign in

Paper Citation Record · LEDGER

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2604.11297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11297 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:34:31.715954Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact25
  • verified fuzzy20
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ea33fd6-dc3d-403b-ba4f-6e4cd7ea5b54 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:16:05.102223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:0484ef05e364debfc5c07d459eb81cc4c8e4240e826bd0171d532ef161ebb9cd

Observation dde50c55-7418-4b3a-a646-fa3c4f5f4ab2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T15:35:32.831908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:d796d6fee71b818ae727413d05cadd14c479d0541fd3854559a06f93061a9b9e

Observation b21c6ace-e1f4-4553-8d50-6b21cb945071 · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.826796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:0026597709c4d9914e3bd4590a91560819fd00ba0bc0468ac6f68dc30a111b39

Observation c2d69158-f2b0-4ad5-94c3-01f4f57f2a5b · outbound

This paper cites Execution-basedcodegenerationusingdeep reinforcement learning.Trans.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Execution-basedcodegenerationusingdeep reinforcement learning.Trans

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.460744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:2ea52a16c96ce9d6587a288412cac53421bc530ba268c3783e0262ccab0764ac

Observation d640a3bd-07c7-4972-badd-a7b0503d3907 · outbound

This paper cites Christiano, Jan Leike, and Ryan Lowe.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Christiano, Jan Leike, and Ryan Lowe

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.464559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:ecbdd0165d6a1ad5128d08d763680523764b0c8356d2acf0e2d97e2184a41aed

Observation 46665d13-9645-4259-a5b0-37c611dc06d0 · outbound

This paper cites arXiv preprint arXiv:2503.01067 , year=.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping arXiv preprint arXiv:2503.01067 , year=

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.816007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:2b18586a0d311cc2117a79c02abd3f4a6b0d3653ec6bb973b0561c3e2f2abbb4

Observation 1dd6fcf5-e8da-4ec2-9e1e-f8ca91e0f7d9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:35:32.820950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:7016af17bb56088b005a38d1e69d7c86204e6d19131da8b62a9bcb24bd722db1

Observation 7aa33035-2676-4431-b5d7-fe5e91bf9747 · outbound

This paper cites Expected return causes outcome-level mode collapse in re- inforcement learning and how to fix it with inverse probability scaling.CoRR, abs/2601.21669.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Expected return causes outcome-level mode collapse in re- inforcement learning and how to fix it with inverse probability scaling.CoRR, abs/2601.21669

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.801473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:c584d51efa6798f89ae8bbd0bcfb83426a3eff24213de077188451b54d1edf24

Observation 7d7b1b52-dcec-47ca-a772-3eab1aab6467 · outbound

This paper cites EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:35:32.836910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:b9de50129e603cd881b5f6042324b97b32c0303ce426b85c1adbc05bff48e270

Observation 9b09873b-c376-45a9-a2e5-0add5efe78e1 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.795743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:8d5af26d4a3d33c1462479b1e40286351581785fa32cb11f88923e9d0efeef01

Observation 92cfaa51-ace3-4881-8f16-8b0abdee0d69 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.446227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:a131e747f646ba0d38df06cfd8a6243c4e8a61e016b517ab400f152ae54a2ac3

Observation 1796ad6a-0103-4283-90e8-518143746bb9 · outbound

This paper cites DSDR: Dual-scale diversity regularization for exploration in LLM reasoning.arXiv preprint arXiv:2602.19895.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping DSDR: Dual-scale diversity regularization for exploration in LLM reasoning.arXiv preprint arXiv:2602.19895

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:05.095573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:b93d41caf1bf652ea67c8dc1f1b3aa4f7654d807a11c8d855e647ac0a87553c6

Observation 50586630-192a-4d37-b726-4078121ca7b1 · outbound

This paper cites Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.438417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:26b67d68fc01f863eeac59c52753736408169eb966cac16b81108bf50830ba01

Observation 9e15d907-0407-4af1-8e92-1c9507d29ac0 · outbound

This paper cites The neural basis of human error processing: Reinforcement learning, dopamine, and the error-related negativity.Psychological Review, 109:679–709, 11 2002.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The neural basis of human error processing: Reinforcement learning, dopamine, and the error-related negativity.Psychological Review, 109:679–709, 11 2002

Reference 14

Resolution
verified exact
doi, observed 2026-05-10T15:35:32.743232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:d599226c6e71dd321584c1cb4c9818ee5190dd4831d8bcb988a6d26ea5019b2f

Observation 59b6c280-e1b0-4950-9013-890dcc8272ee · outbound

This paper cites The Journal of Open Source Software 2(11) (mar 2017).

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The Journal of Open Source Software 2(11) (mar 2017)

Reference 15

Resolution
metadata mismatch
doi, observed 2026-05-10T15:35:32.734582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:790f6d667e2afe45915fb02b3eed5adedc75f8a4704e3cbd9c3e8c4a8a8d526e

Observation 72752f53-1af9-4691-bfa3-3f3127f297b4 · outbound

This paper cites OpenAI o1 System Card.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping OpenAI o1 System Card

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:35:32.723851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:a053551ac1f959e6d338d318fb38fb6622eab65c8f4307243d6b5635819f011b

Observation 9ccf4e04-f567-40fb-898c-0cad0c1e04c7 · outbound

This paper cites Qimeng-codev-r1: Reasoning-enhanced verilog generation.arXiv preprint arXiv:2505.24183.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Qimeng-codev-r1: Reasoning-enhanced verilog generation.arXiv preprint arXiv:2505.24183

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:05.056443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:b73f7511ac385873ea07e3b419768e61da318a574be07426550eeb00e0f180b4

Observation 70c70560-a7fa-4a35-906f-cfabf5963d6e · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.730359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:b9f983ba0334c1213f9b5dfed50dff1256eca9a9eeae10da21c83721170ffe46

Observation 2910a4f1-355f-4a94-8ebf-9d767a333652 · outbound

This paper cites Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:05.083724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:232b3f2d6e319b9594a441a1c85f6f178c2858d6b47d839a956f4b915e0ed00f

Observation ea46ed91-eb05-49a5-a9f7-660f2acd8fd1 · outbound

This paper cites REARANK: reasoning re-ranking agent via reinforcement learning.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping REARANK: reasoning re-ranking agent via reinforcement learning

Reference 20

Resolution
verified exact
doi, observed 2026-05-10T15:35:32.748330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:eee1690c1d2c67d62cad4c5028d34f0a566979587164f560ee98cac1a4bb6440

Observation 3995b01b-fa87-4716-b40c-8a51fc70c80a · outbound

This paper cites Let’s verify step by step.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Let’s verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.396700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:6f25162a6992df7d60b5536d51017d218ad2b7c710887cccd316d35f556283e3

Observation 29894fa3-4f98-402c-bbf6-9272ddeb74cf · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Solving math word problems with process- and outcome-based feedback

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T15:35:32.753463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:e97320d3050c072873a79f31a1e729da16d1f374d92e3f95eb3827840c37ebc7

Observation 7c326e03-d676-43c8-b538-85efe56b9e3f · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Proximal Policy Optimization Algorithms

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:16:05.065074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:a0882556d928091df276d39c30095711eb8cdaed354a45953ede63dbb237caf0

Observation c099b7f6-9a2e-453e-adf0-ac095ad01702 · outbound

This paper cites Rewarding the rare: Uniqueness-aware rl for creative problem solving in llms.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Rewarding the rare: Uniqueness-aware rl for creative problem solving in llms

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.739251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:a3a5166678e9e91d8f373f92bbf86a488a96f4beca9a67ad02a3060e8764050e

Observation 7f61e66d-647e-4ac4-9bb8-3c2b623fc48a · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:05.109966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:e79ea6d00c6e44a733ae33e1b48afa12ae1d0ac85b77c26e2606c7fd441c0ee4

Observation ec353b52-5182-4790-b97c-e9927c819038 · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.810462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:1089fa993ee9223a4c965d79a6afb221267a14b894b4d42ed1f48178643a001a

Observation 176ffacb-3cee-4ccf-8105-48c887858eba · outbound

This paper cites Emogen: Emotional image content generation with text-to-image diffusion models.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Emogen: Emotional image content generation with text-to-image diffusion models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.854791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:9d0ce80bb43d185ecca3bcc4a404c2347f23dc69910efab73528fea2a769f2d9

Observation 80a1572b-eebf-4928-8bc6-9f159a55171c · outbound

This paper cites Multi-objective evolution of heuristic usinglargelanguagemodel.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Multi-objective evolution of heuristic usinglargelanguagemodel

Reference 28

Resolution
verified exact
doi, observed 2026-05-10T15:35:32.805054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:105a9d60dd57ae40ee27fff6823061c92f322bdc5897585e1cf0c8ff367f05d7

Observation a949e1ae-6f48-48d1-9a7e-922d7fcce932 · outbound

This paper cites Latent reward: Llm-empowered credit assignment in episodic reinforcement learning.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Latent reward: Llm-empowered credit assignment in episodic reinforcement learning

Reference 29

Resolution
verified exact
doi, observed 2026-05-10T15:35:32.840809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:70f8b92aa372b1561e838c0881b5bfe701eaf52565481cb467d25b0f6d9f2e5d

Observation c0beafab-6a79-4067-87b4-80c40e6eec9d · outbound

This paper cites Revolve: Reward evolution with large language models using human feedback.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Revolve: Reward evolution with large language models using human feedback

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.434285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:4561ef88d4d7d6921b4ba3da935af410a4a0ff27dd68e0543249d6655f26d93e

Observation 31af2722-e6e6-499a-808b-eaf668f1bd25 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Eureka: Human-level reward design via coding large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.417631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:182d776fdd8b7aad02eb108ac15654b8b98a5cb151bb93fbf16af37f7786d44f

Observation 2ae3172b-5fe5-41e2-93fd-551c5931cbc1 · outbound

This paper cites an unresolved cited work.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:55:11.404832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:c1baa6c5f324c6bcfae937eaf28de5a3f5fc1446c7687e8a596dfb04bc41a0ef

Observation ebfd1f85-9597-4a85-ac28-acc07ae2d4b6 · outbound

This paper cites Locating and editing factual associations in GPT.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Locating and editing factual associations in GPT

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.407839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:79c13d06f7286f26fe61b3ae277d8bb32d9673628a15c0bae0266830182cc226

Observation 59b915ce-fee4-4436-8fb4-f90d7f577115 · outbound

This paper cites Daniel Freeman, Theodore R.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Daniel Freeman, Theodore R

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.410802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:3729df6bf834ac05697de5d5195d28a47ebb5a97b5992872ed5087b1543ceccd

Observation 0b0d8499-8447-4727-8b51-1ba97563fe5e · outbound

This paper cites In-context learning and induction heads.Transformer Circuits Thread.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping In-context learning and induction heads.Transformer Circuits Thread

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.402128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:cae1bec0eafea272da8abccd0ab5fa8a43b87cba0b43be97ef2ad2bd1ad4aacd

Observation c39b9499-16e6-4950-b3df-ba9b0e5c3b20 · outbound

This paper cites Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.846244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:da61c0424600b1da8ec1123a90b3d6f01afd26a8f4bc43cf09158e41a9f17b50

Observation 01801f0c-9b33-4699-b891-51ecca4002d3 · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.765091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:df1ddf4f9e6afe3fb6787e21aec0b3bc9c1feccaf145dbc44571a0a483cb0e29

Observation 7c865a2e-4829-4c33-9139-64133ac26ff9 · outbound

This paper cites Verifyingchain-of-thought reasoning via its computational graph.CoRR, abs/2510.09312.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Verifyingchain-of-thought reasoning via its computational graph.CoRR, abs/2510.09312

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.759051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:358e641faa54b6c7dc2dd34a393f7b597bfe6a949befa764bca8cfcd4e820042

Observation 74c4eec3-f643-4301-b029-6ecfe88d43e0 · outbound

This paper cites Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:23.552990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:51eb25fc249febdda0069423e1b79c1d665065fcc0abb27722c7f6658452aa13

Observation 1f4da862-42a4-4883-92f6-6af5516f5b65 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:54:49.681471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:661905dbfa9787d36ff7c50ffeeee59dc790f32d906a7eae191352520d0f2795

Observation 519e5c6a-547b-473c-88ef-f169412f3fb6 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Measuring mathematical problem solving with the MATH dataset

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.413835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:5a4233b06a1a0a4e16fa1906e1400fd697cf2916490b7719d341f4c898ebbb00

Observation 17af816a-4cc0-47c4-ba6a-27210a2467d3 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Understanding R1-Zero-Like Training: A Critical Perspective

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.785086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:ba5504d2be999c744868b32d2c411f8e7321e4ad937de37b3a5916e1f6e25599

Observation f26a4de9-4d78-4f9b-809c-ff1a8440c331 · outbound

This paper cites Hybridflow: A flexible and efficient RLHF framework.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Hybridflow: A flexible and efficient RLHF framework

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.421900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:e0c4f0158242a3cb24b7d8b58f394596421aa4a22a5aac1d390b5b5823c3b47b

Observation bcac0cb1-e92d-4dbc-aed3-164c6c91d074 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Hybridflow: A flexible and efficient rlhf framework

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.790653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:a6a3d433387206474b7d79e1549bb38359b02c175d60f4b0afd4a1a3ce4efaf1

Observation de4032af-9a3c-410e-b58e-aabe62879214 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Reasoning with Exploration: An Entropy Perspective

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:20:56.524904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:93992cc84723a0da9383474ddc4bc60d9563fcc19f5d7082bd3afa6d438843e3

Observation f09b3a45-9878-4089-908f-0326eac4c8e8 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13(9):9.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13(9):9

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.429684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:50eb1e619e200e821f87215d9683145cf1eb2d813033e3932b44b97d87df9ecc

Observation f9515d54-e8fe-48c1-9bbd-5606c72efdf5 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur- Ari, and Vedant Misra.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur- Ari, and Vedant Misra

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.442217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:80ed58309054ce162d1dcb8fa8880e319a9582ae8073d45dee5cce4732d75b7c

Observation 9d979fce-46d0-43f8-8d39-f56e3b96a242 · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 48

Resolution
verified exact
doi, observed 2026-05-10T15:35:32.779371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:ac324101166aa48e439e714f09a937e225ff00de7529ed425808a92a71a50c74

Observation c4e071d3-2b33-455d-ae8e-8d1518f2d3f1 · outbound

This paper cites 1 𝐾 𝐾Õ 𝑖=1 min 𝑟𝑖(𝜃)𝐴𝑖 ,clip(𝑟𝑖(𝜃),1−𝜖,1+𝜖)𝐴𝑖 !# −𝛽𝔻KL[𝜋𝜃∥𝜋ref] (1) ℒDAPO(𝜃)=𝔼𝑞∼𝒟.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping 1 𝐾 𝐾Õ 𝑖=1 min 𝑟𝑖(𝜃)𝐴𝑖 ,clip(𝑟𝑖(𝜃),1−𝜖,1+𝜖)𝐴𝑖 !# −𝛽𝔻KL[𝜋𝜃∥𝜋ref] (1) ℒDAPO(𝜃)=𝔼𝑞∼𝒟

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.381815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:fc45b203cfc3e54ff1449c3a5f4dc819d458f17bbf2337eaec8ce72382885358

Observation c1a15149-0883-499e-970d-7a5a053e60ee · outbound

This paper cites Since89 2 is closer to 8085, we use it: √ 8085≈89.9166, and thus: 𝑝= −1+89.9166 2 ≈88.9166 2 ≈44.4583, which isn’t an integer.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Since89 2 is closer to 8085, we use it: √ 8085≈89.9166, and thus: 𝑝= −1+89.9166 2 ≈88.9166 2 ≈44.4583, which isn’t an integer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.388731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:8e407864e71d848564daf4714c0203ad16c7d0c4a8b2548c27fec2d408c88413

Observation b9f9fc8c-6c00-424e-8b7e-46cdb7af9e72 · outbound

This paper cites The factors of𝑝2 are1, 𝑝, and𝑝 2.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The factors of𝑝2 are1, 𝑝, and𝑝 2

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.391549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:2635b3c7f5cb93a4d7093164f5eb3f22b20ec08467fee75d7c3c57e7f8fed785

Observation ad69b496-dc57-4420-9395-118cc600c70b · outbound

This paper cites "" Returns a sorted list of all divisors of n.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping "" Returns a sorted list of all divisors of n

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.385106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:47e4f934a2f2e02a59b04132f24d4f91f27f86a2d8844fe2ef69f32b7bfa094d

Observation 9c47dc01-abe8-490c-90d7-7053fcedcbcb · outbound

This paper cites The sum of these three divisors is1+𝑑+𝑛 𝑑 =2022.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping The sum of these three divisors is1+𝑑+𝑛 𝑑 =2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.394049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:f7a9fe95073ead6d0ca59544d30d582859f2e426c6832385f0451fa2d68dc439

Observation 5b250739-dda4-4d81-8bfe-5265199745fd · outbound

This paper cites reasoning path.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping reasoning path

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:55:11.399405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:2cb34acd6024a1557f5b1fe9d203ccfd9153b924c07336beaed6b8f620e863af

Pith citing papers

No inbound Pith citation observations are available.