Pith. sign in

Paper Citation Record · LEDGER

A Survey of Direct Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.11701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11701 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:44.129023Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1c935f6-3a28-48d5-aa9a-b50a8efa91e9 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A Survey of Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.129023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.129023Z digest=sha256:929a234aeb5cdc360a3f491c764502f6644d0290a8fce9b0094de095ba8f08ec

Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · inbound

Intra-Trajectory Consistency for Reward Modeling cites this paper.

Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:34.053645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:34.053645Z digest=sha256:6ff1d2b2261d9e1965f9149223d54fd8a877679a4499f2f675d995cae3940907

Observation 7a998fd7-a7b3-4929-9f05-f782f3d09b1b · inbound

A Novel Self-Evolution Framework for Large Language Models cites this paper.

A Novel Self-Evolution Framework for Large Language Models A Survey of Direct Preference Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:33.764607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:33.764607Z digest=sha256:ec559aed64db67bfff89e27b52ee1804c72f9df6b3127cf5892c904ae38fd8cb

Observation 8de070ee-5aae-42de-9bcb-a29fff768523 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment A Survey of Direct Preference Optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.437263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:d6e2e52af0e0df475529e578be63d3ea67e9710124cf1a837c2183b144127456

Observation ecd42d58-be7e-4676-946a-a1142d2fc2ae · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.231474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.231474Z digest=sha256:3a0637183eeed403208e0eb3b1731c7aa711bf3c7bdf4e44903951d58152960b

Observation ee5a3204-80e7-427e-b045-6d032292e33a · inbound

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models cites this paper.

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:54.350069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T09:38:32.977517Z digest=sha256:c62813379eca53effcf970b26d4806313d3a6f2b11b9f62df606f4be619c858f

Observation 1696a3ff-944e-435f-8c6c-275e5452166a · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey of Direct Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.684688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:5f8a04c1f866218d167ca75cbaa94972e15ee559896016b0aca3be24685c460c

Observation fd377ff2-8a5f-44e1-bfc0-9fec693a5a7f · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization A Survey of Direct Preference Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.255353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:b30a26351f440c814b7e298151ad7fada7c2047968253ff352b803d60b1d404c

Observation 9c1d35da-0d6c-45d9-8dd0-cd8bd70ac6c5 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:16.671074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:abdafde2805c6ced6fdeccc26eee68bba0b2927c25ab735969c6c3f8976bb263

Observation 6b978201-41a7-494c-915b-f4ea47e82fac · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.018076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:8ae1864f215c633284f4ee923704a2b79e0b96bf31086b4ba6d3c66587a9cc3b

Observation 62e96638-a393-4a34-84f2-c538b10ce10c · inbound

TUX: Measuring Human--AI Tacit Understanding cites this paper.

TUX: Measuring Human--AI Tacit Understanding A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T21:22:38.748817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T21:18:37.366258Z digest=sha256:b1aa0013cac627b6fb10c58eed25b3c832f021362320d70754e6ab612568c089

Observation 20ac0933-1d06-433f-bea2-65123a72cac4 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.939868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:37244474219af5812478c42ecf7aaf27c2a4927c0eb30b86e7a688ef20b6a26b

Observation 35783b72-aa3d-4470-8121-0eaf885c9031 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text A Survey of Direct Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:56.375550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:56.375550Z digest=sha256:eb4ec7cba428a4e8c91dd693f6bd9110e3e9f6b711875bb1ed08f683ec2b9386

Observation 6f2b483e-719d-4b54-a9a6-702675fde444 · inbound

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing cites this paper.

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T00:22:51.122740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:22:51.122740Z digest=sha256:3ccf2dcd3676f92d8c9dc4bc65b280e70c2a5a4027047490806ad7ee834aa91b