Pith. sign in

Paper Citation Record · LEDGER

A Survey of Direct Preference Optimization

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.11701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11701 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.171743Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b99ca145-d560-4b7b-9ab4-46d5e33770e3 · inbound

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm cites this paper.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm A Survey of Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.171743Z digest=sha256:82882dd0ad6674d4860d08d4d8ea45942198d3dea6be21de451778e0d3747960

Observation e1c935f6-3a28-48d5-aa9a-b50a8efa91e9 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A Survey of Direct Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.129023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.129023Z digest=sha256:d317847174e11e32b182d9a2070d23d2d52ef6ad114aa1542ec310df01394ec2

Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · inbound

Intra-Trajectory Consistency for Reward Modeling cites this paper.

Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:34.053645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:34.053645Z digest=sha256:62c1dca60ee781a341f386a07143fded4ff83ac88eca9d43fbc9d5be21091975

Observation 7a998fd7-a7b3-4929-9f05-f782f3d09b1b · inbound

A Novel Self-Evolution Framework for Large Language Models cites this paper.

A Novel Self-Evolution Framework for Large Language Models A Survey of Direct Preference Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:33.764607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:33.764607Z digest=sha256:e601f5311ec209b1672a250beeaade99e190e1427748366fee86f05dee9e05ea

Observation 8de070ee-5aae-42de-9bcb-a29fff768523 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment A Survey of Direct Preference Optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.437263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:3bfc69432edda87ccd898ff6b50b12001928a6e2b9dbd758dd9842f382ae9431

Observation ecd42d58-be7e-4676-946a-a1142d2fc2ae · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.231474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.231474Z digest=sha256:862d960a72f0aa80e9675110e2b4a7de604cb0d9cf6e62a36c9c2839baa72e2c

Observation ee5a3204-80e7-427e-b045-6d032292e33a · inbound

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models cites this paper.

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:54.350069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T09:38:32.977517Z digest=sha256:eb920a63b15d3bfa007a92a4538dbc4646800be4a4cd214454e8435d8626cf1f

Observation 1696a3ff-944e-435f-8c6c-275e5452166a · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey of Direct Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.684688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:cee5774ade1c256ab7bf9635eb3d3e8b6198fa92020914db59f78eb8892b205e

Observation fd377ff2-8a5f-44e1-bfc0-9fec693a5a7f · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization A Survey of Direct Preference Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.255353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:bcafb8ed06f42f6bafa470d07409277aa261c03f0efb3f3f4bd4a7e0144adffb

Observation 9c1d35da-0d6c-45d9-8dd0-cd8bd70ac6c5 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:16.671074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:fc987f6f9afd38e44840f183d00252881f1d7a292246337a230535dc6ded4765

Observation 6b978201-41a7-494c-915b-f4ea47e82fac · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.018076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:8c0b53a42408c54b7d74970cd74e85f67ece5cd313e8c399a154f452a686d09b

Observation 62e96638-a393-4a34-84f2-c538b10ce10c · inbound

TUX: Measuring Human--AI Tacit Understanding cites this paper.

TUX: Measuring Human--AI Tacit Understanding A Survey of Direct Preference Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T21:22:38.748817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T21:18:37.366258Z digest=sha256:63cc78d9a90103bbb199190afcba2c111c3a9e0a88c49314ad457e417d2c05b6

Observation 20ac0933-1d06-433f-bea2-65123a72cac4 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs A Survey of Direct Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.939868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:ae01cc9a7d015c9b9cdefca2045b1038418d55665ce70343e49cfd4de85d2832

Observation 35783b72-aa3d-4470-8121-0eaf885c9031 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text A Survey of Direct Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:56.375550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:56.375550Z digest=sha256:93d7bbcf8c42f8ae5707be926b8b0d8e88ceec281caec96c16b53fbd71aa9510

Observation 6f2b483e-719d-4b54-a9a6-702675fde444 · inbound

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing cites this paper.

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing A Survey of Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T00:22:51.122740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:22:51.122740Z digest=sha256:03ecbc2d940cd4d2a5903da6e6d782f3f3f5d036db46a67c5d0d3ca93dda5870