Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2502.19613.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:41:48.404875Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 97f62b28-d998-499f-aa58-9eb0923c1a72 · inbound
Scaling Test-time Compute for LLM Agents Self-rewarding correction for mathematical reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1bd704-6917-4eda-a7cf-32a282e1a413 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-rewarding correction for mathematical reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06743f5c-b56f-4ab0-a5ae-fcede9ccce86 · inbound
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? Self-rewarding correction for mathematical reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5830c2c4-5a28-44d0-9fea-33295906121a · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-rewarding correction for mathematical reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b896a909-8106-4d06-ac64-1436ef0624b9 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f621f4-3604-428e-8cc2-e0bf2c72abc3 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dbfc3b72-9643-43e4-85a6-9b411c13cfd7 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e07ccf-0df1-4722-b143-7e48f81aada3 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-rewarding correction for mathematical reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c7cb4b3-0047-45bd-a6ed-f07c5cd35174 · inbound
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Self-rewarding correction for mathematical reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf61d227-58ea-4435-9385-6261c2420a7c · inbound
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f414539e-896d-4e51-bb05-6e65dc61007e · inbound
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dd7aacd-e04f-4881-a417-eea4ab73c190 · inbound
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning Self-rewarding correction for mathematical reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03fe794f-2c38-44d0-961d-2ae2f83409a2 · inbound
Trust Region On-Policy Distillation Self-rewarding correction for mathematical reasoning
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9f7a4fa-0632-4437-b449-b34408d22a14 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Self-rewarding correction for mathematical reasoning
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86fff995-f14c-4448-b088-6cff9d7586e5 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation deeadaf6-7404-490b-a02c-e0291d1527f5 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cf7f342-4619-40d3-8df7-e5cc461b49d7 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-rewarding correction for mathematical reasoning
Reference 238
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af7e3537-efcf-4951-ba7a-d7583bc4b8f7 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8428ec32-f496-4353-9265-0eb5a8b61781 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.