Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.12854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12854 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:50:38.778526Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:33.817155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation beacce03-b4a2-4061-b7a1-b2e2731349e0 · inbound

Mathesis: Towards Formal Theorem Proving from Natural Languages cites this paper.

Mathesis: Towards Formal Theorem Proving from Natural Languages Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:38.778526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:38.778526Z digest=sha256:ce3a58abc770461ae61a73b17388dd6c5467b7372ed37d0849a90362d2b71c5f

Observation 450bebd0-af38-4ece-ac96-ebf204566c01 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.417741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.417741Z digest=sha256:4deb6a1ca210d930cd27fa8a8f88bce7c53987a5e92426582e3755fb6d3bd141

Observation e2d8a1c9-c418-4448-bdfa-652d2a27dfff · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:41:54.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:e8af4425ab4631ede0d0416b901a4722ad939e413c56e6d4ebb04c0312c9eb41

Observation 62801cd4-ce5a-4f1e-abfb-8c526f5297b7 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:40.888746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:40.888746Z digest=sha256:adb02f340ff5964213d0b394334a53b59e558e620b592a6daf36f5d0fdf3f473

Observation 921189dd-55a7-47c7-8c07-88391aa3e3ff · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:49.468251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:49.468251Z digest=sha256:89c4b37d83304bdb03721ac2b365edb6e359519d41cc5334a28ad446b2ba9d63

Observation ae29fb2c-7c16-4ba5-b2d9-2e85298977f4 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.124148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.124148Z digest=sha256:4e339fb7e6fc3633c1261c8f52be9ef097f0e2c3fd375bb9f520196141f5dc2f

Observation aa1fc4f7-5128-4fc3-9ee6-e1f725596d84 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.074027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:c276628735e3b9b7eab4f7d7a0c50881d3c066250ca8c1fd3e04bee843aa0626

Observation 41a09afb-bd6c-4ae4-bfab-0ba069dbc723 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:25.953549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:65be6ba46866da53dd0d5d25a25905023ca53b0dfd40967f855baa44939d4104

Observation fd439143-50f9-432f-86e7-ec30bd916e06 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.895403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:b489bc4e2a843467476a14bef655d75398b1b03120d49c270c85d0d11c92d60b

Observation 8c53353d-3041-44a5-9005-e0d8f5f1affb · inbound

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR cites this paper.

How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:24:00.555652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-21T06:22:55.286281Z digest=sha256:eb2ed74e9e7d2fe404582272fb077e439c70dd8e00884a829da4d710c1cc4837

Observation 173be850-c0a6-47df-9987-16f2e2589fb1 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.528307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:8431ab669a8b9739a1bafddaaf69ff74e88bec03e96ccc15abdccee53693974b

Observation a6ce62f4-2f0c-435d-9727-edc4d7977589 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:33.820741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:f55c0fb045d7943618ec16eae31d27ce9ba345e3f2a6724802bfa75f26713c51