Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:40.426831Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 15 inbound Pith citation observations for arXiv:2505.24726.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:40.426831Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:41:42.259895Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.690995Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b9ff29d-d9cf-4137-85b9-8e79fd5e86f0 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c883d1e-40c5-4807-8656-7b743d85b4b3 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2dcd2d-6c02-4f85-9673-068e45be6863 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956def62-0880-44ac-8938-04c0604a95dc · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3023ca1a-1ca2-48b4-bb0b-26130eaf3afd · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a66e9ab-d617-413c-829c-94811bb8051b · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f1c250-c8ee-4e31-b444-9d6d0efdd326 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c8d135-793a-4674-80d6-dc6505968bc6 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ReZero: Enhancing LLM search ability by trying one-more-time
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5889e0bd-b0e3-4df6-9ed9-b9dcd554087d · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53714535-dfd3-4a37-967c-cf26cc185c26 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b6c8ef-a99b-4405-a181-c033fad9e685 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee72f43-7db4-4f1a-a9df-fdfd65ad5220 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Distilling the Knowledge in a Neural Network
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b10b17-ef42-4ad3-8070-76876adfe558 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09642e74-6af0-4d90-8c7c-a9d1a5ee1eac · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7878b40-7727-4454-98d5-0f3700fc303f · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53328d92-8663-4c04-81a8-1470c472d2c5 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2e5de6-cd90-47f5-b1d6-7060f75b29d5 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey on Large Language Models for Code Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d92912-ce03-4e3f-8e69-d21f4706a7bb · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1541ad01-913f-4259-a889-09f58ef76593 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Understanding Catastrophic Forgetting in Language Models via Implicit Inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe7c305-c796-4728-8633-132075ffb7a0 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbfd8cb-edec-4ca0-8209-690b52d2dfc7 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b09f2b-dbd3-4689-bafb-059b08b030d2 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3576f0f6-14d8-4d0b-b505-566183825cff · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9997fed6-014f-4fa3-862a-65087ae21702 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11ae6f74-d7e9-4c7b-a813-6b023325a1f7 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c946d2e-a0a2-42c5-aa15-3cf6328d92dd · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97e5b161-4742-443a-ae01-c04c694f00b6 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa751dfe-6a10-4d89-a3c1-a93ec7d09798 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983b2f28-0bfc-438c-865e-fa5d93f2dacd · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75b95565-6a73-42ce-ba8a-5f485c757876 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7cecc57-05b5-467a-9235-e4fba1d354e0 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23a9c6e-29a7-488d-bae5-93536a9ba9a9 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bccfb7d-e822-47f3-92ed-c01937437e1c · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275de774-9fe7-4dc6-af7f-7a46307ec865 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9942cffa-1630-492b-ad08-7bf151dd11b5 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bd7e0c-e0c2-415c-a0dc-b67d52b03bb1 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77215a42-c208-4c9c-9949-fae72e537a42 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee18af6-9c36-4fb0-8713-0e60c8d392b2 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02844c99-a114-41f7-9d76-a1a08b570bf3 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3709454-5482-4e16-bde9-04e5fbf14a8e · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 475ad1e2-c136-4ea0-b489-30e396a9f81a · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a3a9574-7702-447c-a809-de5cd05f0aa2 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Rethinking Chain-of-Thought from the Perspective of Self-Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb982c97-2962-40af-9879-54a35a1185f1 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7abd710-2915-4279-ab23-728769dfb765 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Qwen2.5 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b04294-cced-44c2-8898-8f13660ec3aa · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96df79b1-acab-4ed4-830e-2ec77e7bd15d · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3e9429-4f9e-46dd-87b8-9618be2330e3 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7da77fb-c34e-461d-a598-28986c3a01e3 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4318b2-588c-4f1c-a06d-79e89cc600a5 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning A Survey of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993c7cac-ac4c-4875-8537-9b21c8d774d0 · outbound
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning Xing, Hao Zhang, Joseph E
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a667597f-34d6-4b4d-9ba0-6732f76f3c75 · inbound
A Survey of Context Engineering for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ce0801e-0b98-4eb3-a5ea-bb705d2d323a · inbound
Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cebbe30e-d707-4075-8231-559d79d3907a · inbound
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be380109-018d-4b57-8c52-e97617f69d9d · inbound
Self-Reflective Generation at Test Time Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6e74c5-d46e-4cb1-afb7-3b1f56757a43 · inbound
Agentic Reasoning for Large Language Models Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afb29695-d9ba-4290-944c-d6dcef50184a · inbound
AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b762b6cf-007e-4468-9520-1324bb89998d · inbound
AIPO: Learning to Reason from Active Interaction Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6678781-0785-4e35-aee2-0c6a59a5bc0b · inbound
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2731606-a626-4501-b535-cf4aa6885698 · inbound
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d07f7d49-830e-40ee-9429-65ec56ffab8a · inbound
Trust Region On-Policy Distillation Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 158
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd02e5db-fb1c-4588-b3fd-7b9a18e0a419 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7134ada1-c2e2-46d9-9fa5-bf9de84df25e · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959d28d3-e55a-4719-a13c-ac9a3d57164b · inbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2f4927-bba8-4ab2-b334-9ce615378880 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 782637b5-80bb-422a-9834-55f9aa088fb6 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 151
Source-reported events for the cited work
Unavailable: canonical work link unavailable.