Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2503.12854.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:33:49.468251Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T19:02:33.817155Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e2d8a1c9-c418-4448-bdfa-652d2a27dfff · inbound
Sample-efficient LLM Optimization with Reset Replay Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 921189dd-55a7-47c7-8c07-88391aa3e3ff · inbound
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae29fb2c-7c16-4ba5-b2d9-2e85298977f4 · inbound
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1fc4f7-5128-4fc3-9ee6-e1f725596d84 · inbound
IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 41a09afb-bd6c-4ae4-bfab-0ba069dbc723 · inbound
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd439143-50f9-432f-86e7-ec30bd916e06 · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8c53353d-3041-44a5-9005-e0d8f5f1affb · inbound
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 173be850-c0a6-47df-9987-16f2e2589fb1 · inbound
TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a6ce62f4-2f0c-435d-9727-edc4d7977589 · inbound
Enhancing LLM Metacognition via Cognitive Pairwise Training Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.