Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2409.02392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:58.073304Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:29:15.139325Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e5c32733-42ce-466c-8ebe-df0b96b1af6c · inbound
Training Language Models to Self-Correct via Reinforcement Learning Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a663a407-f076-41a0-aa8d-4a3156ce8dff · inbound
Self-Challenging Language Model Agents Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2ad5c0-c823-4ec9-b278-869e749f1dc5 · inbound
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bad5619-96cf-4268-b912-10a1a64f4afe · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b04d36b0-1656-40f2-8d63-6d43fe8c1702 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ee25b76-35cb-4735-947c-3663b87fcf69 · inbound
REVES: REvision and VErification--Augmented Training for Test-Time Scaling Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64ad0b9e-d5ce-4ea7-91cd-68ce8925bf6f · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8355aac3-cb93-42bb-ab12-945ce6089aec · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Building Math Agents with Multi-Turn Iterative Preference Learning
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.