REVIEW 12 cited by
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Direct Preference Optimization (DPO), which derives reward signals directly from pairwise preference data, has shown its effectiveness on aligning Large Language Models (LLMs) with human preferences. Despite its widespread use across various tasks, DPO has been criticized for its sensitivity to the SFT's effectiveness and its hindrance to the learning capacity towards human-preferred responses, leading to less satisfactory performance. To overcome those limitations, the theoretical understanding of DPO are indispensable but still lacking. To this end, we take a step towards theoretically analyzing and understanding the limitations of DPO. Specifically, we provide an analytical framework using the field theory to analyze the optimization process of DPO. By analyzing the gradient vector field of the DPO loss function, we find that the DPO loss function decreases the probability of producing human dispreferred data at a faster rate than it increases the probability of producing preferred data. This provides theoretical insights for understanding the limitations of DPO discovered in the related research experiments, thereby setting the foundation for its improvement.
Forward citations
Cited by 12 Pith papers
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...
-
SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents
A history-aware guard model with an LLM judge is reported to cut jailbreak success on mobile agent tasks from 86.1% to 8.4% while keeping task completion unchanged at 77.8%.
-
Explicit Preference Optimization: No Need for an Implicit Reward Model
EXPO is a pair of explicit preference-optimization losses that provably avoid DPO's uniform-regularization and poor-interpolation failure modes and outperform DPO on Anthropic HH and IMDb.
-
Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists
HiTEC improves LLM tool calling by embedding hierarchical error checklists in prompts or using them to generate negative examples for KTO fine-tuning.
-
Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy
A three-agent domain-independent framework with distribution-balanced DPO training reaches Combined 106.3 on MultiWOZ 2.2 with Qwen2.5-7B, the best score among the compared baselines.
-
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
A pipeline that generates executable programs from domain knowledge documents and uses them with extracted variables to solve domain-specific calculation problems, improving accuracy over baselines in legal and medical QA.
-
SPRec: Self-Play to Debias LLM-based Recommendation
SPRec alternates supervised fine-tuning with preference optimization, using the model's own prior recommendations as negative examples, to reduce popularity-driven over-recommendation in LLM-based recommenders.
-
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
A contrastive token-scoring method identifies 'critical tokens' in incorrect reasoning traces and penalizes them in DPO, yielding small but consistent accuracy gains on math benchmarks.
-
Continual SFT Matches Multimodal RLHF with Negative Supervision
nSFT matches multimodal RLHF performance by converting rejected responses into corrective SFT data, without pairwise preference optimization.
-
Salamandra Technical Report
Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.
-
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Cal-DPO modifies DPO by adding squared losses that anchor the policy's implicit rewards of chosen and rejected responses to plus or minus 1/(2*beta), reporting gains on reasoning, summarization, and dialogue benchmarks.
-
A Survey on Large Language Models for Mathematical Reasoning
Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.
Discussion (0). Continue with ORCID to comment.