REVIEW 10 cited by
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Direct Preference Optimization (DPO) improves the alignment of large language models (LLMs) with human values by training directly on human preference datasets, eliminating the need for reward models. However, due to the presence of cross-domain human preferences, direct continual training can lead to catastrophic forgetting, limiting DPO's performance and efficiency. Inspired by intraspecific competition driving species evolution, we propose a Online Fast-Slow chasing DPO (OFS-DPO) for preference alignment, simulating competition through fast and slow chasing among models to facilitate rapid adaptation. Specifically, we first derive the regret upper bound for online learning, validating our motivation with a min-max optimization pattern. Based on this, we introduce two identical modules using Low-rank Adaptive (LoRA) with different optimization speeds to simulate intraspecific competition, and propose a new regularization term to guide their learning. To further mitigate catastrophic forgetting in cross-domain scenarios, we extend the OFS-DPO with LoRA modules combination strategy, resulting in the Cross domain Online Fast-Slow chasing DPO (COFS-DPO). This method leverages linear combinations of fast modules parameters from different task domains, fully utilizing historical information to achive continual value alignment. Experimental results show that OFS-DPO outperforms DPO in in-domain alignment, while COFS-DPO excels in cross-domain continual learning scenarios.
Forward citations
Cited by 10 Pith papers
-
Test-Time Scaling via Error Localization
TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.
-
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
Phi-Ground models achieve state-of-the-art click accuracy on five GUI grounding benchmarks for models under 10B parameters using a 40M-sample training recipe with text-first inputs, random-resize augmentation, uniform...
-
Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
Applying DPO with preferences from people with intellectual disabilities improves German simplified-text readability but reduces semantic fidelity, and inconsistent target-group preferences prevent statistically signi...
-
Do not Abstain! Identify and Solve the Uncertainty
The authors build ConfuseBench for three sources of LLM uncertainty and show that asking a follow-up question and checking whether its answer is unique improves source identification.
-
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Sampling dispreferred completions proportionally to the current reward model, as in contrastive divergence, improves preference-optimization performance and is framed as NLL estimation.
-
Understanding the Logic of Direct Preference Alignment through Logic
Direct preference alignment losses can be expressed as logical programs over model predictions, yielding an organized landscape of billions of definable losses and a route to new variants.
-
Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis
Gene-R1 combines domain pre-training, GPT-o1 distilled reasoning examples, and GRPO reinforcement learning so lightweight Llama models match commercial LLMs on gene set analysis.
-
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
DGRO decouples the KL regularization coefficient in reward optimization into two hyperparameters and shows strong reasoning results, though its reward-variance ablation is confounded.
-
Direct Advantage Regression: Aligning LLMs with Online AI Reward
DAR aligns LLMs through advantage-weighted supervised fine-tuning on online AI scalar rewards with dual KL regularization, and reports win-rate gains over online RLHF and online preference methods.
-
Salamandra Technical Report
Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.
Discussion (0). Continue with ORCID to comment.