Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2410.15595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.15595 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:26:01.675314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 519dfa99-5557-425f-8788-51f51b6b75ce · inbound

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts cites this paper.

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T08:00:12.781392Z digest=sha256:a291a1f6549d32b12990971a3115e7a41985a7df60e1ea628d9d939e7c801307

Observation 9e4f5b25-ea62-4f09-bbf0-0c4acd5eda14 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:39371d279824c5ad4adaf30338181529630f2f7e9a3b6c83ac837cd1f6127d6f

Observation 53a7ab6b-16a0-4246-bd5b-83e320dda34c · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:49d2e0017f1fb387965abad1c68a4b7c41ee577184892c660f54cceef468bbc4

Observation 98c858db-038a-4c6e-9a94-6d9fc951bdf2 · inbound

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization cites this paper.

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T05:26:01.675314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:26:01.675314Z digest=sha256:08975b40d7c1cd11405d43aa6fe29c4d00fb23d03c068c54f99b83e91ae8fcab

Observation 3013383f-6f4a-4153-95b5-e41680157898 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:337048076241b9daed110161b72e782928ca6f10ed83c64e39fe23b9a6a54ee3

Observation 72b588ce-8bf1-44d0-8e97-596c09337a06 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:a0b60540584d200bcbad90ad68326fb7f4a832340492a3c22a4573739875e385

Observation 942de73f-1e2a-400a-883e-7b98baa62f04 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:0403bf84f07f212feafa810b514fb3548e745d8baf57afd4e5d767c9269fb37a

Observation 02f081b9-9386-4801-9343-faad5808d4b8 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:b7aa296d6b82ad2238151e339b65cc7c91346739f921fd2308d04e652a9302d0

Observation 5d75473b-9ac6-4a3d-b4b1-9e20570849a5 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:4a0d798e44f66767995aced2c06867a930c7c65f7926b4af22d73563d82e6f83

Observation bbb00dc8-427e-4690-98d6-7d2477d6ca53 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:ab85e1fa1be5621ceea1bde7c6a3d633cbd8f60cc9677affbf3ffe017e70994e

Observation 88539dbc-8626-4e97-ba06-9cded6ea9dac · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:2ac53085224df16514b835001ed102ce222f3014d4418bac1316d7cc9bbd7abc

Observation d2a656bc-9798-4349-a83e-5b4be92d3abf · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.439296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:1c69fc6c2631d369ef8119fb9bf2a29171f0d1ba999d7e6555f2df9001848154

Observation ab5e5094-07ec-4c79-9063-00f45ac2f4b2 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.431552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:15:53.639048Z digest=sha256:82d957eecc25430629736f9244f0e5319196965319d6be27fbc3c81521ca801a

Observation 97fe8aa6-09c7-463d-80d0-c2b16d2dc714 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:19:02.875664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:1fcaf2511ab3a1af123fcef3f422d905fe8f1ce2f6b089cf087237662c65a68b

Observation e86b1d7e-f9db-4e1d-b288-cefa5e6e7f8a · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 233

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.699362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:822aeeb02b2ee79e1d933fe095b76ff19a5baec24f7936cc4d34eb72d1b40592

Observation 65e0b408-5a26-4e3f-ae87-b9879ef5963c · inbound

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design cites this paper.

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:18.707247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:18.707247Z digest=sha256:c043747bc8cb08e28c171c66795697ca9d4aa259b5497caff32e195c5e3f2d91