Pith. sign in

Paper Citation Record · LEDGER

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 15 inbound Pith citation observations for arXiv:2505.24726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24726 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:40.426831Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:41:42.259895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.690995Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b9ff29d-d9cf-4137-85b9-8e79fd5e86f0 · outbound

This paper cites online" 'onlinestring :=.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:34.861649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:34.861649Z digest=sha256:34489af18329dccedf5f80b210a626e4f297563536cc30d24100c66407afcabf

Observation 5c883d1e-40c5-4807-8656-7b743d85b4b3 · outbound

This paper cites write newline.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:34.960301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:34.960301Z digest=sha256:78d063baf6c7f2b8cbafaf1ad08284903ffb6af852b5cf0fec8670ab0c1a60b1

Observation 3b2dcd2d-6c02-4f85-9673-068e45be6863 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.083696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.083696Z digest=sha256:046f88cc5860a6d7dd853e9426c4d6e440eccae331b34956ec4da548aa644e51

Observation 956def62-0880-44ac-8938-04c0604a95dc · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.259812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.259812Z digest=sha256:82100b3ddda20e4a6f72368800aa50e1c98257b90560d9dd5e4fcbf9adce8d36

Observation 3023ca1a-1ca2-48b4-bb0b-26130eaf3afd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T12:35:40.840220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:35.409519Z digest=sha256:265e10c59f2eb4f6c9c79558322949637385db9df099f6d7c2f014dbfd9ebab3

Observation 1a66e9ab-d617-413c-829c-94811bb8051b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.568647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.568647Z digest=sha256:eb316f60ef4fa4e864086112cf110036cfe59cf4271f0e62db2728904894735c

Observation 63f1c250-c8ee-4e31-b444-9d6d0efdd326 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.832113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.832113Z digest=sha256:79ca0616342e4eaf7831f1af8277cbad6d9f90b6fe2dfb068bfbf359c18762a2

Observation 48c8d135-793a-4674-80d6-dc6505968bc6 · outbound

This paper cites ReZero: Enhancing LLM search ability by trying one-more-time.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ReZero: Enhancing LLM search ability by trying one-more-time

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.992263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.992263Z digest=sha256:74577a440bc83a3f80c36442e3c2ab03bb402eb1b2ffe106faa684473eaeee7e

Observation 5889e0bd-b0e3-4df6-9ed9-b9dcd554087d · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.205203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.205203Z digest=sha256:b6d7313e3d3758a1d619826bc3032af00d5ceb87a6ef83971811a1fbfc9167c0

Observation 53714535-dfd3-4a37-967c-cf26cc185c26 · outbound

This paper cites The Llama 3 Herd of Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.351366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.351366Z digest=sha256:887a54005d4e142783b70d3c1b3340e7497a04036d70114f9e68c49d75bc25ff

Observation 34b6c8ef-a99b-4405-a181-c033fad9e685 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.415581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.415581Z digest=sha256:204aef2a62df2c15e0a498b60da5202523b4a2bf847b9b5a36f82b697e840d1f

Observation 1ee72f43-7db4-4f1a-a9df-fdfd65ad5220 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Distilling the Knowledge in a Neural Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.557807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.557807Z digest=sha256:babe773b4312d6ae4953e213f8435b0c76cd3ee9717cadbdb0923c3e9b642e7d

Observation 01b10b17-ef42-4ad3-8070-76876adfe558 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.636735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.636735Z digest=sha256:06ac9635dc0f88b28fb6f7589868a3f438aca21a43cc45902e127f1a1fa7716e

Observation 09642e74-6af0-4d90-8c7c-a9d1a5ee1eac · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.738019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.738019Z digest=sha256:c477c0fa7f9ef0d135cde8549ee07d1e57afbf4d8d4dca76499a26458dcc4080

Observation c7878b40-7727-4454-98d5-0f3700fc303f · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:43.087654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:36.863086Z digest=sha256:30b18028fc16e7daf590184bdc48a3e9662ebf760457fd749f14da5cbd49fcc0

Observation 53328d92-8663-4c04-81a8-1470c472d2c5 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.957944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:36.957944Z digest=sha256:0c9d1726d79f57335bda67a4e402fa0be7c4b56d907f45c1bd4bbba53397dcf4

Observation 9e2e5de6-cd90-47f5-b1d6-7060f75b29d5 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey on Large Language Models for Code Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.056537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.056537Z digest=sha256:87afba558acd3f1434eeaf5e10599f485c65eead860cd0d70807eed51c974d44

Observation 04d92912-ce03-4e3f-8e69-d21f4706a7bb · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.174034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.174034Z digest=sha256:9edf1ae74d9796f478a0067af7a117b8c4caba38c005291b66b753b63cad8c09

Observation 1541ad01-913f-4259-a889-09f58ef76593 · outbound

This paper cites Understanding Catastrophic Forgetting in Language Models via Implicit Inference.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Understanding Catastrophic Forgetting in Language Models via Implicit Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.237089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.237089Z digest=sha256:96592ae82030455b136b2ee9b7a68b8ba5746d02ce9bca580abd35aad50826dd

Observation 7fe7c305-c796-4728-8633-132075ffb7a0 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.311069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.311069Z digest=sha256:f287a0abc255c4da8ce32daac4e85df118ddbbf097399b134171bf520d4213b7

Observation 5cbfd8cb-edec-4ca0-8209-690b52d2dfc7 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.411678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.411678Z digest=sha256:fb4d65ceb6ee7f6b12ce66b5736f8b482964d7ec36c3aca0b1aae7efc9e4a21c

Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.503027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.503027Z digest=sha256:8bf8d3228d2b8b784db268386c9ccc67b6fca27c53f027becb9696bee072cd76

Observation d4b09f2b-dbd3-4689-bafb-059b08b030d2 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.897758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.598472Z digest=sha256:dce15db7d7fcf19a794c560af6e164cf4c1c1252831bf379e1c5cb92632c6aaa

Observation 3576f0f6-14d8-4d0b-b505-566183825cff · outbound

This paper cites Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.669914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.669914Z digest=sha256:ea69e79ec4b14c9e07abbe738bc7205ce32ddca57dc2364987c1819646285e1f

Observation 9997fed6-014f-4fa3-862a-65087ae21702 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.649406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.756243Z digest=sha256:267c51ea3dfd5d473a9d7f2f556fe1bfd3f4af292a9cd1bc2a57bb6b19c4d8ec

Observation 11ae6f74-d7e9-4c7b-a813-6b023325a1f7 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.400357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.873250Z digest=sha256:b1a8a90f92941f4fa4ed81b203d589e457754eb11ff437331928336bb0935492

Observation 7c946d2e-a0a2-42c5-aa15-3cf6328d92dd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.207703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:37.976934Z digest=sha256:795ba97866122fc118313522da622862132405caef2f90ff3e52adbdc87e592b

Observation 97e5b161-4742-443a-ae01-c04c694f00b6 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.088525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.088525Z digest=sha256:b0c3bd725dcb5b314d383fa7aa17a29f84404e5bfb0a1b53968cb04e09f1dc16

Observation aa751dfe-6a10-4d89-a3c1-a93ec7d09798 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.172948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.172948Z digest=sha256:3f026503c0fbe1f4dad0c02725c7494d50ee62b31ff2d2faafdbeb57b8c91c95

Observation 983b2f28-0bfc-438c-865e-fa5d93f2dacd · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:42.055686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:38.271932Z digest=sha256:e482ee9af3c2ea74eb9e260d4d7b922b15dff5fb0ada6d98146b9785736ef4a6

Observation 75b95565-6a73-42ce-ba8a-5f485c757876 · outbound

This paper cites LEMMA: Learning from Errors for MatheMatical Advancement in LLMs.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning LEMMA: Learning from Errors for MatheMatical Advancement in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.363254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.363254Z digest=sha256:a793bc36344162f6edf2b9bab93f91ffda93ed417dba91d872b6e2f61432d598

Observation c7cecc57-05b5-467a-9235-e4fba1d354e0 · outbound

This paper cites Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.436062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.436062Z digest=sha256:8436b1532d629ea1496b5c67164905f06958ee9d88e19cd8c56b4d8301ed36e6

Observation c23a9c6e-29a7-488d-bae5-93536a9ba9a9 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.568666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.568666Z digest=sha256:6773962ef3d1de6bfe0287d5f271406ccc8d4a57c5335281b72c7e0a0e19ec0f

Observation 8bccfb7d-e822-47f3-92ed-c01937437e1c · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.675580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.675580Z digest=sha256:799be84837913589a4311417f4ff937117e406322e80e41b8ba90523e056bf71

Observation 275de774-9fe7-4dc6-af7f-7a46307ec865 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.889507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.889507Z digest=sha256:481b64e80027bac975fc1b0769edc4bfe5432b39a288d960b8ed9f933c080983

Observation 9942cffa-1630-492b-ad08-7bf151dd11b5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.941943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.941943Z digest=sha256:432a88783c9c7b9acad787f9ecee0b3feacd01fc3eac41da11fb503c09137a01

Observation d2bd7e0c-e0c2-415c-a0dc-b67d52b03bb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.013972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.013972Z digest=sha256:9c345d2ddf37b3cd18d5e79f1b4380e94cb47c658dd2a0b9157b91a80e371dcc

Observation 77215a42-c208-4c9c-9949-fae72e537a42 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.089985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.089985Z digest=sha256:c04bf3865e590282e51368ee1196ab25bd7065c5eab1003c827589d1d70ffc9c

Observation 3ee18af6-9c36-4fb0-8713-0e60c8d392b2 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.164960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.164960Z digest=sha256:5c2d6d0f868c6df23b21f35c6180246d586c6fa9a638a2a024c9ee26200d521c

Observation 02844c99-a114-41f7-9d76-a1a08b570bf3 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.242501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.242501Z digest=sha256:febe587c2bf92586d78a452fad5956173934e9dfd718f2f5c931753cc3e3501c

Observation b3709454-5482-4e16-bde9-04e5fbf14a8e · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:41.839084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:39.320511Z digest=sha256:d2eed8eea5a57a0b9954e77c4676128645dee728e98508f791c1b5520e86859c

Observation 475ad1e2-c136-4ea0-b489-30e396a9f81a · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:41.639978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:39.427431Z digest=sha256:d25d82080fb9b443c612e0ad3d1b88ed3ace2ac38850e1c8de63717d1e1f3ad4

Observation 7a3a9574-7702-447c-a809-de5cd05f0aa2 · outbound

This paper cites Rethinking Chain-of-Thought from the Perspective of Self-Training.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Rethinking Chain-of-Thought from the Perspective of Self-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.496054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.496054Z digest=sha256:c5795bb1cd62d6898265958feb9201a17db75d98671be736dccc60a10b243b42

Observation bb982c97-2962-40af-9879-54a35a1185f1 · outbound

This paper cites Qwen2 Technical Report.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.602927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.602927Z digest=sha256:d904eee39007529404f9f2e0fb5c9621c4b9d1d8eccae0f6f28307999ec139ae

Observation c7abd710-2915-4279-ab23-728769dfb765 · outbound

This paper cites Qwen2.5 Technical Report.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.687345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.687345Z digest=sha256:c5d3f93980fad9b081b29a00e4a3d59befa618eb68decad17d6310befb63acad

Observation a2b04294-cced-44c2-8898-8f13660ec3aa · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.783549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.783549Z digest=sha256:dc9cce59f96d927af8b536e6d4563e5f2ff35e7dc6430cfcf87843fc471ba286

Observation 96df79b1-acab-4ed4-830e-2ec77e7bd15d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.880812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.880812Z digest=sha256:5e7613c726af09bb8de7f975baf2df2b1d1b3be292e3bbc5ac41931431d95cd5

Observation 6d3e9429-4f9e-46dd-87b8-9618be2330e3 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:39.948340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:39.948340Z digest=sha256:e8206436d88a17cdd70dd8633e5c9505a446df95eb4c5f8a0278d88182163990

Observation b7da77fb-c34e-461d-a598-28986c3a01e3 · outbound

This paper cites an unresolved cited work.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.097436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.097436Z digest=sha256:edcee3820d200fc69fb4d6c44fa1bc049bf2c1eec419b936ef2ac645ee9515b4

Observation 3a4318b2-588c-4f1c-a06d-79e89cc600a5 · outbound

This paper cites A Survey of Large Language Models.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.252367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.252367Z digest=sha256:b36b81fc64bfcded517e2c196f4638030c12f040a57e4703d0d06902a688f579

Observation 993c7cac-ac4c-4875-8537-9b21c8d774d0 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Xing, Hao Zhang, Joseph E

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:40.426831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:40.426831Z digest=sha256:3d48e77a0b5a729e59c159eb2ffe9690fb7e8e6e88d2663d3250e9671c1aff22

Pith citing papers

Observation a667597f-34d6-4b4d-9ba0-6732f76f3c75 · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:58:45.141925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:abd618a10f6c021cac035fe86f45f38b1a623785d93554fbea1c3d1665538d86

Observation 4ce0801e-0b98-4eb3-a5ea-bb705d2d323a · inbound

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency cites this paper.

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:46:53.907850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T22:42:53.663354Z digest=sha256:537285a6c76720f852e46671675ec0e3c8d51ea5a38419e012c965b9550c95d3

Observation cebbe30e-d707-4075-8231-559d79d3907a · inbound

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts cites this paper.

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:42:38.721025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:42:07.883909Z digest=sha256:fc369ed09b162ad8d1b52e47de17753aef5ca0bbcb8a60cd10281a232bd75f41

Observation be380109-018d-4b57-8c52-e97617f69d9d · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:42.259895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:42.259895Z digest=sha256:9a5848fe02d504b55e263f9e306b1bf94185e1a89d0c90f6c3495eac745c0144

Observation 2c6e74c5-d46e-4cb1-afb7-3b1f56757a43 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.410357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:9c3b419ce1d93820fbd485abe63e15232899f73f2d8d6140bf1ae092cf5fb7a3

Observation afb29695-d9ba-4290-944c-d6dcef50184a · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.794258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:c620e7fade2b4afbbc2d106cb40c8ad9e38483ec9790339d27e563e9e9e8d4a0

Observation b762b6cf-007e-4468-9520-1324bb89998d · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.262051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:769ba2ed00d3e64430f22169ff640587f70b648aabec4400a0286e9c786ae4b0

Observation d6678781-0785-4e35-aee2-0c6a59a5bc0b · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.376226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:40:23.895430Z digest=sha256:4a6e60413a4996f829efbde1ae42634631a3301d9bfa4a65942c4f03a009a079

Observation b2731606-a626-4501-b535-cf4aa6885698 · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.314539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:48:49.322744Z digest=sha256:2e5f57a3cc9e7ad7e4b76f343bddd6142385e35c8aa86ea1a29b1c97a178b0d4

Observation d07f7d49-830e-40ee-9429-65ec56ffab8a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.720223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:8c706861b11b42981086c21e2c486c9a71adbe2d5cf6c003b55c0f777c714cbf

Observation dd02e5db-fb1c-4588-b3fd-7b9a18e0a419 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.453282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:ff482ca1f0fd9485efc2d78efe06dfcead3df649aa00bc3054f0f4076937846d

Observation 7134ada1-c2e2-46d9-9fa5-bf9de84df25e · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:16.219670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:16.219670Z digest=sha256:c266a7597494591359523bd444cedc04a5a5dd0647add1c05bdc67bcd57fe4eb

Observation 959d28d3-e55a-4719-a13c-ac9a3d57164b · inbound

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry cites this paper.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:53a0988c5c33796d190abf54aada12e1d0d8a5a40c868e493ec39f2f83202299

Observation cd2f4927-bba8-4ab2-b334-9ce615378880 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 151

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.692261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:0bc3fa13e21cda35c90ca4d93af7f2f374a8c10b9898dc273e11d88073919610

Observation 782637b5-80bb-422a-9834-55f9aa088fb6 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:18af4fb17d66c4773f84fa7055083a2b5c041ab37db1a3af3ed1436d3395c771