Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2503.12854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12854 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:33:49.468251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:33.817155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e2d8a1c9-c418-4448-bdfa-652d2a27dfff · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:41:54.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:f67e16310e84a75c34949fa92ae021a5d688ed31a42175d3844ced8c36357485

Observation 921189dd-55a7-47c7-8c07-88391aa3e3ff · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:49.468251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:49.468251Z digest=sha256:9d8f99563a03c7555b32f064bbdb9f666ebeade985df22049f59afd05fb08608

Observation ae29fb2c-7c16-4ba5-b2d9-2e85298977f4 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.124148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.124148Z digest=sha256:dfbff881700b51679970aee3e386d41852b7b28925a17880dd9f799d94c0ae28

Observation aa1fc4f7-5128-4fc3-9ee6-e1f725596d84 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.074027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:cfd196ac66bb9d76f2bf75ddee30f401a9239c5f3a6831fb2b74eefd6848b10d

Observation 41a09afb-bd6c-4ae4-bfab-0ba069dbc723 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:25.953549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:4333c4a8ef449d260a045551f51ee6e71f1ffaf28403065641649fbaa092f4f5

Observation fd439143-50f9-432f-86e7-ec30bd916e06 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.895403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:79fd704310adf52155a114755aa93d0bffe1b72a1b9a44a058c15d343521f705

Observation 8c53353d-3041-44a5-9005-e0d8f5f1affb · inbound

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR cites this paper.

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:24:00.555652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-21T06:22:55.286281Z digest=sha256:235e5e5222469f72fb9b0c7da5bcc4cf09582411db97780a7084ba8c3ad911c0

Observation 173be850-c0a6-47df-9987-16f2e2589fb1 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.528307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:5cd4dd96dbd7d3ba0ffd89dbbb514a850099d1f5965bdc355e73be687a8be895

Observation a6ce62f4-2f0c-435d-9727-edc4d7977589 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:33.820741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:2ed05392200210b36128710369f02084a17cdaf953e34679b4762b10a7e1d22e