Pith. sign in

Paper Citation Record · LEDGER

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 15 inbound Pith citation observations for arXiv:2505.24726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24726 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:40.426831Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:41:42.259895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.690995Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b9ff29d-d9cf-4137-85b9-8e79fd5e86f0 · outbound

This paper cites online" 'onlinestring :=.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:34.861649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:34.861649Z digest=sha256:72c00b8f10609a8a7992376342d674f553565dc815d50f05db38dd713fffa061

Observation 5c883d1e-40c5-4807-8656-7b743d85b4b3 · outbound

This paper cites write newline.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:34.960301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:34.960301Z digest=sha256:e1b6e23b14b58ecc175407b45d5d1e0e565f74d66cc0e779cdd324c8f93d10fc

Observation 3b2dcd2d-6c02-4f85-9673-068e45be6863 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.083696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.083696Z digest=sha256:43ef72925af3524727eb92d089f2a85ffd9917c819338b51e1355353b43a34f7

Observation 956def62-0880-44ac-8938-04c0604a95dc · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.259812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.259812Z digest=sha256:d63439dd32fd449005e76bdede2c41f45d5fcf117d70ea17dc82065b81beb9e9

Observation 3023ca1a-1ca2-48b4-bb0b-26130eaf3afd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T12:35:40.840220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:35.409519Z digest=sha256:076c63b5fc8b6bbe15c0072883a2bec682cbb4fa1e32023d516d75cd4156a1ea

Observation 1a66e9ab-d617-413c-829c-94811bb8051b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.568647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.568647Z digest=sha256:bfbca508c0e1e04ee9c500a8b0027d5949d711540e4cc68ec2521f1c08d394f9

Observation 63f1c250-c8ee-4e31-b444-9d6d0efdd326 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.832113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.832113Z digest=sha256:fcbe189c155e04c4606279b72e8e05f32c751ca26942c95345bce797b52aa389

Observation 48c8d135-793a-4674-80d6-dc6505968bc6 · outbound

This paper cites ReZero: Enhancing LLM search ability by trying one-more-time.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ReZero: Enhancing LLM search ability by trying one-more-time

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.992263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.992263Z digest=sha256:9e67fe370c008848c583e6db2e80f24b8701ec31df95ab16fae344979c3e6282

Observation 5889e0bd-b0e3-4df6-9ed9-b9dcd554087d · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.205203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.205203Z digest=sha256:1364c616de9b8ad7c84fabc2be1786159fa3947e1a0e27300504d3ff3e260ca9

Observation 53714535-dfd3-4a37-967c-cf26cc185c26 · outbound

This paper cites The Llama 3 Herd of Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.351366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.351366Z digest=sha256:ee6a7dfdb5804b90ec3ed62257559cc1d5ef7f39f3f1f31f6351b83da0bde9c2

Observation 34b6c8ef-a99b-4405-a181-c033fad9e685 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.415581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.415581Z digest=sha256:ad940ce6da4ce45a08680c021e8d71663b3552e98f1fbcbdb51e42d84ece44be

Observation 1ee72f43-7db4-4f1a-a9df-fdfd65ad5220 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Distilling the Knowledge in a Neural Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.557807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.557807Z digest=sha256:a0189a0d6821836f36cd40ae9bd4a854e83cc0069890e69369bc2fe421198a1f

Observation 01b10b17-ef42-4ad3-8070-76876adfe558 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.636735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.636735Z digest=sha256:7689252eaa5033ef76805becc6e667087e7f57d9400125ab2caf37acbfdf5e16

Observation 09642e74-6af0-4d90-8c7c-a9d1a5ee1eac · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.738019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.738019Z digest=sha256:d3ff829227f5ae8810d582a495949de82f6c8ad43704b7fa876833a7da4c69eb

Observation c7878b40-7727-4454-98d5-0f3700fc303f · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:43.087654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:36.863086Z digest=sha256:d9fbfb7828e11e4d4e4cff8dfd14f9e2a106a12ebbafe1e3b91a465c9f4e20b9

Observation 53328d92-8663-4c04-81a8-1470c472d2c5 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.957944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.957944Z digest=sha256:6dcfffcb05277d0b17e66199c9e872fe23f64dff13c07a29c19237621d4ea17d

Observation 9e2e5de6-cd90-47f5-b1d6-7060f75b29d5 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey on Large Language Models for Code Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.056537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.056537Z digest=sha256:4808b0a10de14d908bf7b311aca288c90fa662959c970f97fb560d3f61bea54f

Observation 04d92912-ce03-4e3f-8e69-d21f4706a7bb · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.174034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.174034Z digest=sha256:c748c5b4e267edfcd42ceb72a4d8cecfe8a1c7fa1666f568b01337039d8b5233

Observation 1541ad01-913f-4259-a889-09f58ef76593 · outbound

This paper cites Understanding Catastrophic Forgetting in Language Models via Implicit Inference.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Understanding Catastrophic Forgetting in Language Models via Implicit Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.237089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.237089Z digest=sha256:d32070e6845c5cb2274a238d311e4a8c75218a3bb9a5095ddce07a3af7bb93e7

Observation 7fe7c305-c796-4728-8633-132075ffb7a0 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.311069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.311069Z digest=sha256:6ea2ea067596709325d191d2fb33928d51542e94494d69e9023608c0cbc52a44

Observation 5cbfd8cb-edec-4ca0-8209-690b52d2dfc7 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.411678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.411678Z digest=sha256:7824fe07ac300d71192ca2aa7ff84132049d2d1cb9e8e9b2dd8a326be38f0caa

Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.503027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.503027Z digest=sha256:7a06a9b2d6b6f8d5ae866d42c76d334b6799d60aa120562da817377f74cb94fb

Observation d4b09f2b-dbd3-4689-bafb-059b08b030d2 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.897758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.598472Z digest=sha256:6a1be3bdbffccbcf06054043b0502dcec58ff0e364a5518377df02362e627b8b

Observation 3576f0f6-14d8-4d0b-b505-566183825cff · outbound

This paper cites Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.669914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.669914Z digest=sha256:207c99643da77ec38fe9c395fb492ae5ef1c6917e550d837fb2e8fcd626a7a20

Observation 9997fed6-014f-4fa3-862a-65087ae21702 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.649406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.756243Z digest=sha256:33c1803a45573c34526cc4afef73e72728d5819d6cce22ad097f4fd4d57df65f

Observation 11ae6f74-d7e9-4c7b-a813-6b023325a1f7 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.400357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.873250Z digest=sha256:a89dc947a95478a763b2f0303084c5a1cafebf9d0a1755e35b4e4bd048bf9ad2

Observation 7c946d2e-a0a2-42c5-aa15-3cf6328d92dd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.207703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.976934Z digest=sha256:3cdbb98a864bf1559a046558b68613f7f7b2c98b9fe6d07fedb5cd074e1b0b31

Observation 97e5b161-4742-443a-ae01-c04c694f00b6 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.088525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.088525Z digest=sha256:d65d9c918ed645837b762f76c9639af6c48c4a7e56858fbac176562c3549b9fd

Observation aa751dfe-6a10-4d89-a3c1-a93ec7d09798 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.172948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.172948Z digest=sha256:fa60eade0b3594314333255ce0bfa823377b865f1bbdeb2382920851517bf781

Observation 983b2f28-0bfc-438c-865e-fa5d93f2dacd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.055686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:38.271932Z digest=sha256:84e3ad17ecc3bfe5cfc4e65fa3bd685a5752ef0bf52f7bb93312b03194b0c531

Observation 75b95565-6a73-42ce-ba8a-5f485c757876 · outbound

This paper cites LEMMA: Learning from Errors for MatheMatical Advancement in LLMs.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning LEMMA: Learning from Errors for MatheMatical Advancement in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.363254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.363254Z digest=sha256:df5226907e4b8673ebadc0c2f19660bba91f2ee47a761c0e7eda5a3dc7d429e4

Observation c7cecc57-05b5-467a-9235-e4fba1d354e0 · outbound

This paper cites Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.436062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.436062Z digest=sha256:9af1d4c304f80d0c26a0aabc284ed5b3f113446fb3118500e51b7e6d58cf57b2

Observation c23a9c6e-29a7-488d-bae5-93536a9ba9a9 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.568666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.568666Z digest=sha256:ee1a6f61f798cb78af85a2f2952aaa18b8fd46e2dd01c41f8160e9d1987bb6a1

Observation 8bccfb7d-e822-47f3-92ed-c01937437e1c · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.675580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.675580Z digest=sha256:4b8c88117230e83f4954adc59d8787dac1709449fa995cc22b82f5fc806920ec

Observation 275de774-9fe7-4dc6-af7f-7a46307ec865 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.889507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.889507Z digest=sha256:a5ace717bacc60133095f72c7fa12d0f9ac461dae3574d0bca3ac2b78fef553f

Observation 9942cffa-1630-492b-ad08-7bf151dd11b5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.941943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.941943Z digest=sha256:1555e41d7a2bbd8a634b5820e54be548d857e0854f339c99f0c11c3b35caa2de

Observation d2bd7e0c-e0c2-415c-a0dc-b67d52b03bb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.013972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.013972Z digest=sha256:aefc331c5d497b6814635e6190716f96e5e6a0219b760ef95b767b8839effc0a

Observation 77215a42-c208-4c9c-9949-fae72e537a42 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.089985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.089985Z digest=sha256:a6ea948175544f302e06806a148247d8e5097f6e4630a39d33df115b87811c28

Observation 3ee18af6-9c36-4fb0-8713-0e60c8d392b2 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.164960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.164960Z digest=sha256:6ac653241fa7e14165ed8d52daf1c17e5d057d6f3661602014c59452094419c2

Observation 02844c99-a114-41f7-9d76-a1a08b570bf3 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.242501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.242501Z digest=sha256:1c8832a3b1072c110d3f48cc42a633334cde1c0b2506ff423b446bcad3314f55

Observation b3709454-5482-4e16-bde9-04e5fbf14a8e · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:41.839084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:39.320511Z digest=sha256:e5335c96eec47f3b702b626d76a905d4ab644cbb4b7b155eb0af48209fef4998

Observation 475ad1e2-c136-4ea0-b489-30e396a9f81a · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:41.639978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T12:35:39.427431Z digest=sha256:e91f411be0887849f806d0e622e6d3befa80e758b9b6adb86f9b2aaecf091333

Observation 7a3a9574-7702-447c-a809-de5cd05f0aa2 · outbound

This paper cites Rethinking Chain-of-Thought from the Perspective of Self-Training.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Rethinking Chain-of-Thought from the Perspective of Self-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.496054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.496054Z digest=sha256:ca08f4e4832c4c8c82aa7998290dd4e5e0f32c28efe796f4dfaa85846b572ded

Observation bb982c97-2962-40af-9879-54a35a1185f1 · outbound

This paper cites Qwen2 Technical Report.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.602927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.602927Z digest=sha256:7ebf916350b479aaa255c198ad3895201eed8ef9fb527c1b05736e878e86dd28

Observation c7abd710-2915-4279-ab23-728769dfb765 · outbound

This paper cites Qwen2.5 Technical Report.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.687345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.687345Z digest=sha256:d7a556a0f5de352ead1da9356acfc790c28dd9715566a7ac49fab8f8245712b8

Observation a2b04294-cced-44c2-8898-8f13660ec3aa · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.783549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.783549Z digest=sha256:bd2935d1e205019f7f9366588bb30fcea101ed65c944ef255885e6ec89b585ac

Observation 96df79b1-acab-4ed4-830e-2ec77e7bd15d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.880812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.880812Z digest=sha256:ed4c1df55ec6df661100a56e7ac818f16c7ce81dc013627a36d96105b5bd95c5

Observation 6d3e9429-4f9e-46dd-87b8-9618be2330e3 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.948340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.948340Z digest=sha256:45ec999d85db2900f1c7fed2f125d37e48e0fdbc9031f37927e298651287a296

Observation b7da77fb-c34e-461d-a598-28986c3a01e3 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.097436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.097436Z digest=sha256:b812f0137e19d5946556da6a4dc540d325939bac60413c39e49a9692c49ea396

Observation 3a4318b2-588c-4f1c-a06d-79e89cc600a5 · outbound

This paper cites A Survey of Large Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.252367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.252367Z digest=sha256:de715fdfa28fc71cfcd82d7e0a3257b876a75e929e237e42e58311995f819206

Observation 993c7cac-ac4c-4875-8537-9b21c8d774d0 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Xing, Hao Zhang, Joseph E

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.426831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.426831Z digest=sha256:3108f79d9f2bb1f02be908a3b89ac67492983e75c9af7c59a51304e131bd0608

Pith citing papers

Observation a667597f-34d6-4b4d-9ba0-6732f76f3c75 · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:58:45.141925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:43d919abaa4d33fbcd339bfd23abbed69b962089a5d4482b1927e2e9e5a6fa4c

Observation 4ce0801e-0b98-4eb3-a5ea-bb705d2d323a · inbound

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency cites this paper.

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:46:53.907850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T22:42:53.663354Z digest=sha256:0c7f7d8f62eb2dd623ba3330a2c1800ecc1cb681313b6c43137f963b83e24615

Observation cebbe30e-d707-4075-8231-559d79d3907a · inbound

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts cites this paper.

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:42:38.721025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T13:42:07.883909Z digest=sha256:7adcf011c657aa20415055e5e475816b753d0ebb0f13e106c5ce0cc17898800d

Observation be380109-018d-4b57-8c52-e97617f69d9d · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:42.259895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:42.259895Z digest=sha256:d58736245cd7fa6d35ee7c1abff162580b78f0af7e3387806189fed8307a0d31

Observation 2c6e74c5-d46e-4cb1-afb7-3b1f56757a43 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.410357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:e96f85e9d714cdd823974eadf21edbda68083249f8f7b3471cf2e20d62fe130f

Observation afb29695-d9ba-4290-944c-d6dcef50184a · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.794258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:39749290295a2fa86a40391f0ded74f94c32e88a51735670556cd588738835b9

Observation b762b6cf-007e-4468-9520-1324bb89998d · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.262051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:caf7e78736c23a998ad274edb3a70313370c1ced3eae430186e536111aee7832

Observation d6678781-0785-4e35-aee2-0c6a59a5bc0b · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.376226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:40:23.895430Z digest=sha256:06c65d017de9f1e69d2c67b714098b9fae84632ae642cd1404aebb038b3572b0

Observation b2731606-a626-4501-b535-cf4aa6885698 · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.314539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T22:48:49.322744Z digest=sha256:58f8789604bbea7ce2c9acab4cd5ba2b7f5b7146a0fe3b73bd3c04d1475fa306

Observation d07f7d49-830e-40ee-9429-65ec56ffab8a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.720223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:602d1ec9dbed1cacf34977e6bde81b5aa1917ffa7afcb2825909015b16b57ec3

Observation dd02e5db-fb1c-4588-b3fd-7b9a18e0a419 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.453282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:cfc87a0ffc4ff456d95a73c5b27e4587f2210d6c4b73ba4c6a614f005bd21b75

Observation 7134ada1-c2e2-46d9-9fa5-bf9de84df25e · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:16.219670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:16.219670Z digest=sha256:d77461288449f95904f7ed0ab24a7eb0021da245e79b9b60ed463947f53c7804

Observation 959d28d3-e55a-4719-a13c-ac9a3d57164b · inbound

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry cites this paper.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:ad48738fce729ce7262410754afc9183b3fb4b7bca1bf0e100c67d0176f6c353

Observation cd2f4927-bba8-4ab2-b334-9ce615378880 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 151

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.692261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:2363b5a6d799b6f916d633eb33705fd3aa8afb69e1be83f274bd5f8d49d6346f

Observation 782637b5-80bb-422a-9834-55f9aa088fb6 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:16d0b8e45ff450b9db6ab1c8a0e22032190ba244dbf7b2347ff18315b14a1efe